AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
AgentWorld finds even the best tested LLM succeeds on only 52 percent of long-horizon multi-agent tasks.
AgentWorld is an open benchmark of 100 human-annotated tasks and 100 augmented variants for long-horizon collaboration among 3–20 LLM agents. Tasks span more than 50 interaction rounds in an MMORPG sandbox and require communication, joint planning, and resource sharing while each agent acts without access to others' internal states. Causal Collaboration Effectiveness traces which actions actually contributed to the outcome. Tests of Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B found a best task-success rate of only 52.0%, with failures from communication breakdowns, role confusion, and lost shared plans.
- 100 annotated tasks plus 100 variants in an MMORPG sandbox
- Tasks need 3–20 agents across more than 50 rounds
- CCE measures how much team effort causally contributed
- Best tested model reached only 52.0% task success
Full article175 words · extracted from huggingface.co · click to collapse
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.31590