arXiv cs.AI / cs.LG / cs.CL·11d agoCoding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead#benchmarks#coding-agents#evaluationAI research1
arXiv cs.CR·16d agoWhy User Studies and Participant Experience Reporting Matter for VR Motion Privacy?#anonymization#behavioral-biometrics#leaderboardsResearch
Hugging Face daily papers·17d agoBenchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation#benchmarks#datasets#evaluation
arXiv cs.AI / cs.LG / cs.CL·19d agoWhat Does an LLM-Agent Leaderboard Rank Actually Compare?#benchmarks#evaluation-methodology#leaderboardsAI research