BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence
BELIEFRAG tracks evidence state for adaptive RAG, matching quality with 35-39% fewer tokens than iterative retrieval.
BELIEFRAG is a closed-loop controller that maintains an explicit evidence state and chooses retrieval, rewriting, verification, answering, stopping, or abstention. On six QA benchmarks with gpt-oss-120b it reaches mean token F1 0.572 using 3.89k tokens, beating fixed iterative retrieval (0.555 F1) with 39% fewer tokens. The pattern transfers to Qwen3-32B (0.552 versus 0.523 F1, 35% fewer tokens). Analysis attributes gains mainly to corrective re-retrieval, with calibrated answerability the strongest operational signal.
- Closed-loop controller tracks sufficiency, reliability, conflict, uncertainty, gaps, and cost.
- With gpt-oss-120b, mean token F1 is 0.572 at 3.89k tokens per question.
- Qwen3-32B reaches 0.552 F1 versus 0.523, using 35% fewer tokens.
- Gains come mainly from corrective re-retrieval, not pruning alone.
Full article203 words · extracted from arxiv.org · click to collapse
Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which action should follow. Existing methods often use such signals as separate triggers, making it difficult to preserve a coherent evidence state across a trajectory; we call this problem evidence-state fragmentation. We introduce BELIEFRAG, a closed-loop controller that updates an explicit state over sufficiency, reliability, conflict, uncertainty, evidence gaps, and acquisition cost, then chooses among retrieval, query rewriting, verification, answering, stopping, and abstention. Across six QA benchmarks with gpt-oss-120b, BELIEFRAG reaches mean token F1 0.572 with 3.89k tokens per question, outperforming fixed iterative retrieval (0.555 F1) while using 39% fewer tokens. The same quality-cost pattern transfers to Qwen3-32B, where BELIEFRAG reaches 0.552 F1 versus 0.523 for iterative retrieval while using 35% fewer tokens. Analysis shows that the main gains come from corrective re-retrieval rather than pruning alone, while several belief dimensions are redundant and calibrated answerability plays the strongest operational role. Calibration improves threshold stability across related evidence sources, although source shift can still invalidate the same decision signal.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.39139