AI agents build 3D scenes from photos but have no idea if they got it right
LEGO-Anything has coding agents write Blender scenes from one photo, but they cannot reliably judge geometric accuracy.
LEGO-Anything, called Image-to-Code, has a coding agent iteratively write and revise an executable Blender program from a single photo so the scene can be edited and queried. LEGO-Bench scores 208 images from 104 simulator scenes and 443 assets on validity, geometric accuracy, and visual similarity. GPT-6 Astra reached 53.4 percent indoors and 39.6 percent outdoors, versus about 15 percent for weaker GPT setups; extra reasoning lifted Astra from 32.3 to 61.8 percent on an office subset. Models judged geometry near chance, while the training-free LEGO-Plugin, which substitutes measurements for self-assessment, improved weaker agents by up to 62.7 percentage points.
- Image-to-Code agents iteratively write, run, and revise Blender code from one photo.
- LEGO-Bench covers 208 images from 104 scenes and 443 registered assets.
- GPT-6 Astra scored 53.4 percent indoors and 39.6 percent outdoors.
- Geometric self-assessment landed near chance, so agents cannot judge their own scenes.
- Training-free LEGO-Plugin uses measurements and boosted weaker agents by up to 62.7 points.
Full article837 words · extracted from the-decoder.com · click to collapse

The approach is called "Image-to-Code." A coding agent receives a single image and writes code for Blender, the widely used 3D software. Rather than generating the scene in one pass, the agent works iteratively: it writes code, runs it, looks at the result, and revises until the scene matches the original.
Because the output is a program, it captures objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like any other piece of code.
Simulator scenes provide the exact ground truth
To measure how well agents perform, the team introduces LEGO-Bench. It contains 208 images from 104 indoor and outdoor scenes and uses 443 registered assets.
Real photos don't provide a precise 3D ground truth to compare against, according to the researchers. Simple synthetic scenes look unrealistic. So LEGO-Bench splits the difference by rendering its images from professionally built simulator scenes. The inputs look natural, while the exact geometry, depth, and object assignments stay hidden and serve as the answer key for automated scoring. Scene complexity can also be ramped up without changing lighting or camera settings.

The benchmark scores each scene on three axes: validity checks whether a usable scene artifact was delivered at all. Reconstruction measures how accurate the visible geometry is. Appearance captures how closely the look matches the original by re-rendering the submitted scene and comparing it pixel by pixel against the reference image.
Agents deliver usable artifacts but struggle with geometry
All six tested GPT configurations delivered a working scene almost every time. Accuracy varied wildly, though. GPT-6 Astra, the best tested model, hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Weaker configurations scored around 15 percent.

The more complex a scene, the more accuracy drops, and outdoor scenes are harder than interiors. When the researchers increased the models' reasoning budget, the GPT-6 variants improved a lot. Astra's score on an office test subset jumped from 32.3 to 61.8 percent.
To understand why, the researchers analyzed the agents' work steps. They found poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment as the most common issues.

That last point is the most telling. When models had to pick which of two versions better matched the original, their geometric judgments landed near or below chance level. An agent basically can't tell whether its own scene has gotten better. The researchers conclude that refinement should rely on concrete measurements, not the agent's own judgment.

That insight led the authors to build LEGO-Plugin, an extension that needs no extra training. It anchors the starting scene in the reference image, swaps the unreliable self-judgment for concrete measurements, and shields correct progress from regressive edits. The plugin improved all six models. Weaker agents saw the biggest gains, with boosts up to 62.7 percent. The already strong top model gained only about two percentage points.

Reconstructed scenes aren't accurate enough yet
Finally, the team tested whether reconstructed scenes could serve as a basis for standard vision tasks. Because each scene is an executable program, object detection, segmentation, and depth estimation can be pulled directly from it.
Without any extra training, the scenes produced usable but unremarkable results across all three tasks. Object detection fared best, with the reconstructed scenes reaching roughly half the performance of the specialized model DINO. For segmentation and depth estimation, the gap to specialized models like SAM 3 and Depth Anything 3 was larger. The authors say executable scene programs from current coding agents show promise but aren't accurate enough. A big gap remains between a working result and a faithful reconstruction.
GPT-6 Astra's commanding lead in LEGO-Bench lines up with other observations. AI researcher Yoav Artzi sees the model as a major leap in spatial understanding and suspects it was trained on large amounts of 3D data like Blender scenes. 3D software makers are already gearing up for these kinds of agents. Unity has released official plugins for Claude Code and Codex. Other approaches skip code entirely and reconstruct scenes directly inside the model, like the Atlas world model from World Labs. Google Deepmind takes yet another route with GenCeption, using a video model for depth estimation and segmentation that matches the performance of specialized models.