CompoWorld: Compositional Environment Scaling for General Agents
CompoWorld composes 448 services to train Qwen3.6-35B-A3B, beating Claude Opus 4.6 on AutomationBench.
CompoWorld expands agent training by composing reusable services rather than generating tasks inside one environment. Coding agents convert tool specifications into verified services, a world model covers tools that cannot be implemented reliably, and random walks over dependency graphs create cross-service tasks. The authors build 448 services exposing 10,130 tools and train Qwen3.6-35B-A3B with 3,000 supervised trajectories and 1,000 reinforcement-learning tasks. The model gains 9.17 points on average across eight benchmarks and surpasses Claude Opus 4.6 on AutomationBench.
- Library has 448 services and 10,130 tools with shared interfaces.
- Dependency-graph random walks create verifiable cross-service tasks.
- Qwen3.6-35B-A3B is trained with 3K SFT trajectories and 1K RL tasks.
- Average improvement is 9.17 points across eight benchmarks.
- It surpasses Claude Opus 4.6 on AutomationBench.
Full article187 words · extracted from huggingface.co · click to collapse
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (CompoWorld), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33665