WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
WorldSolver tests LLM agents on generating physics solvers across 168 tasks from 61 graphics papers.
WorldSolver is a benchmark of 168 physics-simulation tasks drawn from phenomena in 61 classic computer-graphics papers across seven physical domains. Agents must implement the solver inside a fixed scene scaffold and are scored on execution, visual fidelity, and physical plausibility. Frontier agents struggled: GPT-5.6-Sol and Claude-Opus-5 led with overall scores of 48.7% and 46.7%. Code is released on GitHub.
- Benchmark covers 168 tasks from 61 classic graphics papers.
- Tasks span seven physical domains inside fixed simulation scaffolds.
- GPT-5.6-Sol scored 48.7% and Claude-Opus-5 scored 46.7%.
- Executable solvers were difficult; visual and physical correctness were harder.
Full article244 words · extracted from arxiv.org · click to collapse
LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08720