Coding agents can build an editable 3D scene from a single photo by writing and refining code step by step. A new benchmark reveals where the process still breaks down in self assessment and geometric accuracy. A 3D reconstruction from a single photo is most useful when it exists as an executable program you can inspect, edit, and query. That premise drives LEGO-Anything, a project from researchers at the University of Maryland and AWS.
The method is called Image-to-Code. A coding agent receives a single image and writes code for Blender, a widely used 3D software. Instead of generating the scene in one pass, the agent works iteratively. It writes code, runs it, looks at the result, and revises until the scene matches the original. The output is a program capturing objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like regular code.
Simulator scenes provide exact ground truth
To measure agent performance, the team introduced LEGO-Bench. It contains 208 images from 104 indoor and outdoor scenes using 443 registered assets. Real photos lack a precise 3D ground truth for comparison. Simple synthetic scenes look unrealistic. LEGO-Bench resolves this by rendering images from professionally built simulator scenes. The inputs look natural while exact geometry, depth, and object assignments stay hidden to serve as the answer key for automated scoring.
The benchmark scores each scene on three axes. Validity checks whether a usable scene artifact was delivered. Reconstruction measures visible geometry accuracy. Appearance captures how closely the look matches the original by re-rendering the submitted scene and comparing it pixel by pixel against the reference image.
Read nextAI Beats Licensed Accountants On Speed And Accuracy In StudyAgents deliver artifacts but struggle with geometry
All six tested GPT configurations delivered a working scene almost every time. Accuracy varied wildly. GPT-6 Astra, the best tested model, hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Weaker configurations scored around 15 percent. Scene complexity drops accuracy, and outdoor scenes are harder than interiors. Increasing the models to a higher reasoning budget improved the GPT-6 variants significantly, with Astra jumping from 32.3 to 61.8 percent on an office test subset.
Analysis of the work steps showed poor initial attempts, revisions that undid earlier progress, and unreliable self assessment as the most common issues. When models had to pick which of two versions better matched the original, their geometric judgments landed near or below chance level. The researchers conclude that refinement should rely on concrete measurements rather than the agent's own judgment.
Reconstructed scenes lack sufficient accuracy
The authors built LEGO-Plugin, an extension needing no extra training, to address these flaws. It anchors the starting scene in the reference image, swaps unreliable self judgment for concrete measurements, and shields correct progress from regressive edits. The plugin improved all six models, with weaker agents seeing boosts up to 62.7 percent and the top model gaining about two percentage points. Executable scene programs from current coding agents show promise according to the authors, but they are not accurate enough yet, leaving a big gap between a working result and a faithful reconstruction.



