DAGent: Evaluate-then-Grow Planning for Deep Research Agents
DAGent grows research-agent plans from evidence and beats open baselines on BrowseComp-Plus, GAIA, and xbench-DeepSearch.
DAGent is a DAG multi-agent framework that grows a deep-research plan one batch at a time from confidence and uncertainty signals, rather than committing to a full plan and patching it later. A hierarchical context layer shares compact QueryDocs while retaining full traces, and DAGRPO adds topology-conditioned reinforcement-learning credit. At Qwen3-235B-A22B it surpasses the strongest open-source baseline by 5.3, 5.8, and 2.0 points on BrowseComp-Plus, GAIA, and xbench-DeepSearch; the lead replicates across four open backbones and GPT-5 at 327K context. At Qwen3-8B, DAGRPO improves 3.0 average Pass@1 points over same-budget outcome-only GRPO while using fewer tokens, tool calls, and steps.
- Evaluate-then-Grow expands the task DAG using confidence and uncertainty.
- Leads the strongest open baseline by 5.3, 5.8, and 2.0 points.
- Gains hold across four open backbones and GPT-5 at 327K context.
- DAGRPO adds 3.0 Pass@1 points over outcome-only GRPO at Qwen3-8B.
- Evidence-conditioned plans use fewer tokens, tool calls, and steps.
Full article257 words · extracted from arxiv.org · click to collapse
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.39154