MarkTechPost·1d agoPerplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation#perplexity#glm-5.2#self-distillation 4 min
arXiv cs.AI / cs.LG / cs.CL·2d agoExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds#explorationbench#benchmark#scientific-discoveryAI research 2 sources
Hugging Face daily papers·3d agoPUBG Ally: A Conversational Embodied Agent as an AI Teammate#pubg-ally#embodied-agent#conversational-ai
arXiv cs.AI / cs.LG / cs.CL·4d agoMeasuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation#tool-use#evaluation#ollamaAI research 2 sources
arXiv cs.AI / cs.LG / cs.CL·5d agoCritical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use#reinforcement-learning#tool-use#multi-turnAI research
Hugging Face daily papers·5d agoHarness-Zero: Harness Distillation via Agent-as-Harness#harness-zero#agents#harness-distillation 2 sources
Hugging Face daily papers·6d agoVideoGen-Agent: Reinforcing Video Generation Agents#videogen-agent#video-generation#reinforcement-learning
arXiv cs.CR·8d agoCIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents#llm-agents#privacy-leakage#evaluationAI safety & security1
MarkTechPost·8d agoAlibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use#agentic-ai#alibaba#long-context 5 min1
arXiv cs.CR·10d agoASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions#benchmark#data-exposure#evaluationAI safety & security
Hugging Face daily papers·11d agoBI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence#benchmark#business-intelligence#fine-tuning
arXiv cs.AI / cs.LG / cs.CL·15d agoCanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models#curriculum-learning#diffusion-language-models#reasoningAI research1
MarkTechPost·16d agoGoogle Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation#bfcl#dataset-generation#function-calling 4 min1
Hugging Face daily papers·18d agoMetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes#agents#benchmark#fine-tuning1
arXiv cs.AI / cs.LG / cs.CL·18d agoProcedural Graphs: Self-Evolving Execution Structures for LLM Agents#llm-agents#memory#planningAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback#acebench#bfcl#function-callingAI research1
Hugging Face daily papers·19d agoProcedural Graphs: Self-Evolving Execution Structures for LLM Agents#agent-planning#llm-agents#procedural-graphs1
Hugging Face daily papers·19d agoNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness#agentic-post-training#benchmarks#fine-tuning1
Hugging Face daily papers·21d agoAgentic Visual Generation: From Generative Models to Agentic Control#agentic-ai#llm-agents#survey
Hugging Face daily papers·23d agoOccamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work#35b#agentic-ai#co-work-agents
Help Net Security·25d agoAn AI CAPTCHA solver talked itself out of the right answer#benchmark#captcha#computer-vision 3 min1