DAGent: Evaluate-then-Grow Planning for Deep Research Agents
DAGent plans deep-research agent graphs incrementally from evidence confidence, outperforming open baselines on three benchmarks.
DAGent is a DAG-based multi-agent framework that grows a research plan one batch at a time, conditioning each expansion on confidence and uncertainty from completed nodes instead of committing to a full plan first. A hierarchical context layer shares compact QueryDocs while keeping full traces for recall, and DAGRPO adds topology-conditioned credit to executor rollouts. At the Qwen3-235B-A22B scale it beats the strongest open-source baseline by 5.3, 5.8, and 2.0 points on BrowseComp-Plus, GAIA, and xbench-DeepSearch. The lead holds across four open backbones and GPT-5 at 327K context, and at Qwen3-8B DAGRPO adds 3.0 average Pass@1 over same-budget outcome-only GRPO with lower token, tool-call, and step cost.
- Evaluate-then-Grow expands the DAG using node confidence and uncertainty.
- Leads open baselines by 5.3, 5.8, and 2.0 points on three benchmarks.
- Gain holds across four open backbones and GPT-5 at 327K context.
- DAGRPO adds 3.0 Pass@1 at Qwen3-8B versus outcome-only GRPO.
Full article257 words · extracted from huggingface.co · click to collapse
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.39154