arXiv cs.AI / cs.LG / cs.CL·11d agoCoding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead#benchmarks#coding-agents#evaluationAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoExecCritic: Learn to Test, Test to Improve for Coding Agents#automated-testing#coding-agents#gpt-5.6AI research2
Hugging Face daily papers·25d agoBilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems#game-theory#kimi#memory2