4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench tests whether agents can reconstruct dynamic video scenes as executable graphics programs.
4DCodeBench evaluates agents on 4D inverse graphics by asking them to reconstruct dynamic scenes from video as executable graphics programs. Tasks use real-world videos and synthetic scenes spanning deformation, fluid flow, and fracture, and require abstractions such as physical simulation. Benchmarking of frontier models finds that strong static reconstruction does not yet produce reliable reconstruction of complex dynamics. The benchmark is released at github.com/4DCodeBench/4DCodeBench.
- Agents must emit executable graphics programs from video.
- Scenes cover deformation, fluid flow, and fracture.
- Strong static reconstruction does not yield reliable dynamics.
- Benchmark and code are public on GitHub.
Full article124 words · extracted from huggingface.co · click to collapse
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.03715