EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
EvoDuet co-evolves web queries and solutions, raising LLM discovery scores on 21 scientific tasks.
EvoDuet is a bilevel method that co-evolves scientific-discovery solutions and web-search queries while keeping model weights fixed. A retrieval gate decides whether to fetch new documents, reuse stored ones, or proceed without retrieval, and an inner loop ranks documents by predicted solution score. Across 21 tasks with one candidate per iteration, it lifts OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, while Qwen3.5-9B does not improve. Best runs exceed previously reported scores on eight tasks and also help other evolutionary scaffolds.
- A retrieval gate chooses new search, reuse, or no documents each iteration.
- Inner loop refines queries; outer loop generates and scores candidates in parallel.
- GPT-5.6-Luna gain rises from 74.1% to 78.0%; Gemini-3.8-Flash from 61.3% to 82.3%.
- Qwen3.5-9B does not benefit; best runs beat prior scores on eight of 21 tasks.
Full article192 words · extracted from arxiv.org · click to collapse
Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.40340