Hugging Face daily papers·5d agoPACT: From Credit Assignment to Critic Alignment#reinforcement-learning#llm-post-training#actor-critic
Hugging Face daily papers·7d agoOne to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents#swe-agents#reinforcement-learning#distillation1
arXiv cs.AI / cs.LG / cs.CL·9d agoAn Empirical Study of Harness Design for Coding Agents#agent-evaluation#coding-agents#context-managementAI research2
arXiv cs.AI / cs.LG / cs.CL·9d agoAdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair#code-generation#llm-agents#memory-retrievalAI research
Hugging Face daily papers·10d agoAn Empirical Study of Harness Design for Coding Agents#agent-evaluation#coding-agents#context-management2
Hacker News · AI·14d agoReal-SWE: Benchmarking AI models on private, real-world, enterprise codebases#benchmark#coding-agents#enterpriseAI research 11 min1
Hugging Face daily papers·19d agoSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents#benchmarks#coding-agents#evaluation1
arXiv cs.AI / cs.LG / cs.CL·19d agoWhat Does an LLM-Agent Leaderboard Rank Actually Compare?#benchmarks#evaluation-methodology#leaderboardsAI research