RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
RECAST learns to route retrieval and computation to build evidence, beating strong baselines by about 16 points.
RECAST treats evidence construction as a sequential decision process over retrieval and computation rather than fixed similarity search. A lightweight RouterLM selects operations that a frozen CompilerLM turns into executable code, then passes accepted evidence to a frozen AnswerLM. Trained with supervised fine-tuning and group relative policy optimization, it reaches a 75.6% mean success rate across six benchmark families, 15.9 points above the strongest large-model baseline. A trained Qwen3.5-9B RouterLM outperforms a training-free Gemini 3.5 Flash RouterLM by 5.0%, and three held-out benchmarks improve 15.0% on average.
- RECAST derives evidence via retrieval, computation, and synthesized tools.
- RouterLM is trained with SFT then GRPO; other models stay frozen.
- Mean success is 75.6%, 15.9 points above the best large-model baseline.
- Trained Qwen3.5-9B RouterLM beats training-free Gemini 3.5 Flash by 5 points.
- Held-out benchmarks improve 15.0% on average over the strongest baseline.
Full article226 words · extracted from arxiv.org · click to collapse
Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.10507