AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
AREX-2 trains a Qwen3.8-27B agent on long-horizon reflective tasks, lifting research and coding benchmarks.
AREX-2 synthesizes long-horizon improvement trajectories from machine-learning and algorithmic programming tasks to teach reflection and sustained test-time iteration. An agent built on Qwen3.8-27B scores 81.8 on MLE-bench Lite and 70.7 on Frontier-CS. It transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its round budget grows.
- Training uses verifiable ML and programming improvement trajectories.
- Qwen3.8-27B scores 81.8 on MLE-bench Lite and 70.7 on Frontier-CS.
- Transfers to BrowseComp 84.0, HLE 52.6, GAIA 92.2, and DeepSearchQA 93.8.
- Results keep improving as the iteration budget grows.
Full article158 words · extracted from huggingface.co · click to collapse
We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.38288