Adaptive Consistency Graph for Long-Horizon Agents
Adaptive Consistency Graph raises GPT-5.6-luna long-horizon success from 44.5% to 50.2%, led by BrowseComp-Plus.
The paper introduces the Adaptive Consistency Graph to keep long-horizon language-model agents aligned with the original task as evidence, requirements, and execution state drift apart. ACG stores execution evidence and provenance in a persistent graph and builds a temporary requirement-centered view for each decision within a bounded context budget, without replacing the planner or tool executor. In a matched evaluation, GPT-5.6-luna's average success rises from 44.5% with ReAct to 50.2%, with the largest gain on BrowseComp-Plus (73.5% versus 62.4%).
- ACG stores execution evidence and provenance in a persistent graph.
- Each decision gets a temporary requirement-centered view under a bounded context budget.
- It does not replace the base planner or tool executor.
- GPT-5.6-luna success rises from 44.5% with ReAct to 50.2%, and BrowseComp-Plus from 62.4% to 73.5%.
Full article149 words · extracted from huggingface.co · click to collapse
Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, historical evidence, and the current execution state may gradually become disconnected, so later decisions can drift from the original objective. We study this problem by introducing the Adaptive Consistency Graph (ACG) for long-horizon execution. ACG incrementally organizes execution evidence and its provenance in a persistent graph, then constructs a temporary requirement-centered view for each decision under a bounded context budget. Rather than replacing the base agent's planner or tool executor, ACG provides a structured and traceable context view for each decision. In the matched evaluation, ACG improves GPT-5.6-luna's average success from 44.5\% with ReAct to 50.2\%, with the largest gain on BrowseComp-Plus (73.5\% versus 62.4\%). We further analyze trajectory structure and inference cost to characterize this improvement.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.32754