ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Wei-Jung Huang

What Does an LLM-Agent Leaderboard Rank Actually Compare?

infoAI researchimportance 20
AI summary · glm-5.3-flash

A methodological study shows close LLM-agent leaderboard rank gaps on SWE-bench and similar benchmarks often do not support superiority claims.

The paper defines an estimand-aware pairwise procedure for comparing agents, checking common support and applying explicit uncertainty rules and practical margins. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are frequently unresolved, and proxy labels or utility rules can change which system is selected. The authors argue a leaderboard score summarizes a released evaluation but does not by itself justify pairwise superiority conclusions.

  • Pairwise agent comparisons require estimand, common support, and uncertainty rules
  • Close ranks on SWE-bench, AgentRewardBench, and tau2-bench often unresolved
  • Proxy labels and utility rules can flip which system is selected
  • DataAgentBench and Open Agent show limits of coarse public records
Full article160 words · extracted from arxiv.org · click to collapse

An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.07785