QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
QwenGyre trains extreme-long-horizon agents, lifting Qwen 3.8 2.4T by 6 points on NL2RepoBench with up to 1.85x speedups.
QwenGyre is an online reinforcement-learning framework for extreme-long-horizon LLM agents whose rollouts can last hours, include hundreds of environment interactions, and approach 1 million tokens. It reallocates GPUs between rollout and training without interrupting live executions, and a trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths. Scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it improves NL2RepoBench from 52.5% to 58.5% in 48 steps. Reported speedups reach 1.85 times over Colocate and 1.78 times over Async.
- Targets rollouts lasting hours, with hundreds of interactions and nearly 1M tokens.
- Elastically moves GPUs between rollout and training without stopping live executions.
- Reconstructs branching histories, scores partial progress, and drops redundant paths.
- Qwen 3.8 2.4T rises from 52.5% to 58.5% on NL2RepoBench in 48 steps.
- Speedups reach 1.85 times versus Colocate and 1.78 times versus Async.
Full article155 words · extracted from huggingface.co · click to collapse
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% to 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to 1.85times and 1.78times speedups over Colocate and Async, respectively.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33848