arXiv cs.AI / cs.LG / cs.CL·4d agoGrow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents#llm-agents#agent-harness#webarenaAI research
Hugging Face daily papers·6d agoEDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation#tool-calling#synthetic-data#agents1
arXiv cs.AI / cs.LG / cs.CL·8d agoNemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities#speech-to-speech#full-duplex#tool-callingAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoMTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents#asr#benchmark#llm-evaluationAI research
arXiv cs.CR·10d agoClosed-World Resolution Against Tool Hallucination in LLM Agents#agent-security#benchmark#llm-agentsAI safety & security1
Hugging Face trending models·11d agoCactus-Compute/needle3 — new model trending #30 on Hugging Face#edge-ai#on-device#open-weightsModel release 6 min
arXiv cs.AI / cs.LG / cs.CL·11d agoWhat Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity#llm-compression#model-degradation#moeAI research
arXiv cs.AI / cs.LG / cs.CL·22d agoMulti-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe#benchmark#data-synthesis#fine-tuningAI research1