Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads
An agentic RAG budget study finds broader question coverage reduces evaluation error more than extra reads or trajectories.
The paper studies how to split agentic RAG evaluation budgets across questions, search trajectories, and repeated reads using HotpotQA and MuSiQue. At 34.14–34.39 million model tokens, covering more questions lowers standard error by 33% versus five reads and 12.6% versus three trajectories. Nested and question-only forecasts match those allocations within 4.0% and 3.5%, and depth beyond a two-trajectory audit shows no clear forecasting gain. Temperature zero reduces answer disagreement from 14.3% to 3.4% with similar comparison precision.
- Study compares questions, trajectories, and repeated reads on HotpotQA and MuSiQue.
- At about 34M tokens, broader question coverage cuts standard error by 33% versus five reads.
- Archived forecasts predict allocations within 4.0% and 3.5%.
- More questions beat more trajectories at search prices of $0–1 per 1,000 requests.
- Temperature zero drops answer disagreement from 14.3% to 3.4%.
Full article138 words · extracted from huggingface.co · click to collapse
Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers standard error by 33\% versus five reads and 12.6\% versus three trajectories. Archived nested and Q-only forecasts predict these allocations within 4.0\% and 3.5\%, respectively. Depth subsets establish no clear forecasting advantage beyond the two-trajectory audit. One-read variance penalties relative to the fitted optimum at the same token budget are 0--9.9\%, with substantial Pro uncertainty. Under recorded model fees, more questions beat more trajectories at search prices of \$0--1 per 1,000 requests; question-versus-read fee rankings remain unresolved. Temperature zero cuts answer disagreement from 14.3\% to 3.4\% while comparison precision stays similar. \par\medskip\noindentKeywords: Agentic RAG; Evaluation budget; Generalizability theory; Repeated sampling.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.05034