ScholarCatalyst Benchmark Finds Agents Trail Embeddings
ScholarCatalyst shows agentic search trails embeddings at retrieving papers that inspired 207 computer-science projects.
ScholarCatalyst is a retrieval benchmark built from judgments by 184 lead authors of 207 recent computer-science papers who labeled which earlier works did or could have advanced their projects and supplied rationales. Given an initial research question, systems must retrieve those papers from only the literature available when each project began. Agentic search scored 0.42 Recall@20, below 0.48 for embedding retrieval, and an agent using Claude Fable 5.1 reached only 0.51 Recall@20. The authors call for training that captures expert literature intuition. The Hugging Face note of 2026-09-30 and the arXiv listing of 2026-10-01 agree on the sample, task design, and scores.
- 184 lead authors labeled which earlier works did or could have advanced 207 recent computer-science papers and supplied rationales.
- Given an initial research question, systems must retrieve those papers from only the literature available when each project began.
- Agentic search scored 0.42 Recall@20 versus 0.48 for embedding retrieval.
- An agent using Claude Fable 5.1 reached only 0.51 Recall@20.
- Authors call for training that captures expert literature intuition.
- Hugging Face daily papers listed the work on 2026-09-30; an arXiv cs.AI/cs.LG/cs.CL listing followed on 2026-10-01.
Coverage timelineoldest first · each row is one article
- · 2d agoScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Hugging Face daily papers· 52
ScholarCatalyst finds agents no better than embeddings at retrieving papers that inspire new research.
- · 1d agoScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
arXiv cs.AI / cs.LG / cs.CL· 48
ScholarCatalyst asks models to retrieve papers that inspired 207 CS projects; agents fail to beat embedding retrieval.