Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
Agentic retrieval beats dense search by 8.7 nDCG@10 but is far slower and costlier.
The paper studies agentic retrieval that combines LLM reasoning with corpus retrievers in a ReAct loop for complex search tasks. Using the same embedding model, it improves nDCG@10 by 8.7 points over standard dense retrieval and remains competitive on the ViDoRe v3 and BRIGHT leaderboards, including out-of-domain tasks. The gain is expensive: average latency is 107.4 seconds versus 0.67 seconds, and each query consumes about 764.1K input tokens and 5.8K output tokens.
- Agentic ReAct retrieval gains 8.7 nDCG@10 over dense retrieval
- Same pipeline is competitive on ViDoRe v3 and BRIGHT
- Average latency rises from 0.67 seconds to 107.4 seconds
- Each query uses about 764.1K input and 5.8K output tokens
Full article174 words · extracted from huggingface.co · click to collapse
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.05750