AI agents build 3D scenes from photos but have no idea if they got it right
LEGO-Anything has coding agents write Blender scenes from one photo, but they cannot reliably judge geometric accuracy.
LEGO-Anything, called Image-to-Code, has a coding agent iteratively write and revise an executable Blender program from a single photo so the scene can be edited and queried. LEGO-Bench scores 208 images from 104 simulator scenes and 443 assets on validity, geometric accuracy, and visual similarity. GPT-6 Astra reached 53.4 percent indoors and 39.6 percent outdoors, versus about 15 percent for weaker GPT setups; extra reasoning lifted Astra from 32.3 to 61.8 percent on an office subset. Models judged geometry near chance, while the training-free LEGO-Plugin, which substitutes measurements for self-assessment, improved weaker agents by up to 62.7 percentage points.