DAGent: Evaluate-then-Grow Planning for Deep Research Agents
DAGent grows research-agent plans from evidence and beats open baselines on BrowseComp-Plus, GAIA, and xbench-DeepSearch.
DAGent is a DAG multi-agent framework that grows a deep-research plan one batch at a time from confidence and uncertainty signals, rather than committing to a full plan and patching it later. A hierarchical context layer shares compact QueryDocs while retaining full traces, and DAGRPO adds topology-conditioned reinforcement-learning credit. At Qwen3-235B-A22B it surpasses the strongest open-source baseline by 5.3, 5.8, and 2.0 points on BrowseComp-Plus, GAIA, and xbench-DeepSearch; the lead replicates across four open backbones and GPT-5 at 327K context. At Qwen3-8B, DAGRPO improves 3.0 average Pass@1 points over same-budget outcome-only GRPO while using fewer tokens, tool calls, and steps.