Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarkingnew
DualViewEval compresses agent benchmarks by jointly modeling outcome and process signals, achieving 24x-40x compression with only 20 tasks on APEX-Agents and BFCL.
DualViewEval is an agent benchmark compression method that jointly exploits outcome and process relations from trajectories to learn exact-size minisets predicting full-benchmark scores. The authors analyze large-scale trajectories and identify six process signals systematically associated with final agent performance. Across five agent benchmarks and five baselines, it achieves the best results on all datasets: with only 20 tasks it reaches 24x-40x compression on APEX-Agents and BFCL, reduces MAE by 14.5%-28.2% over the strongest competitors, and improves Kendall's tau by up to 7.2% relative to EssenceBench on SWE-bench Verified.