T-Search: An Open Agentic Retriever and Playground for Hard Multi-Step Search
T-Search, fine-tuned from Qwen3.6-35B-A3B, reaches 56.0 Recall@10, 14.4 points above its base.
T-Search is an open-weight agentic retriever that runs bounded multi-round search over a fixed corpus and returns ranked evidence chunks with short justifications, leaving answer generation to a separate model. It is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts. The authors release the model, harness, demo, and three benchmarks, including TRuST, described as the first native-Russian hard-search benchmark.
- Open-weight retriever built on Qwen3.6-35B-A3B.
- Trained with round-sliced SFT, then GSPO on a recall reward.
- One rollout scores 56.0 Recall@10, 14.4 points above the base.
- Three fused rollouts reach 61.3 and beat larger open models.
- Releases weights, harness, demo, and the TRuST Russian benchmark.
Full article129 words · extracted from arxiv.org · click to collapse
We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator can be swapped without retraining. T-Search is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic search tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks with gold evidence annotations, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts, outperforming larger open models. We release the model, harness, live demo, and three benchmarks, including TRuST, the first native-Russian hard-search benchmark.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06782