3DCodeVerse
Exploring 3D Worlds as Code
Yipeng Gao1,*,†Ziyao Zeng2,*Jinfa Huang3,*Yijiang Li4,*
Haolin Xiong1Wang Qin1Xinlei Yu1Youheng Yao2Jie Zhang5Kun-Yu Lin6Rong Liu1Yinyuan Zhao7
Yajie Zhao1Meiqi Guo8Jiebo Luo3Lei Shu8Yunhao Ge1Xinyu Hu9Laurent Itti1
1University of Southern California2Yale University3University of Rochester4University of California, San Diego5Columbia University6Nanyang Technological University7University of California, Santa Cruz8Google DeepMind9Philo Labs
* Equal contribution† Project Lead

The Jade Guardian, running live. Every mesh, material and light, and the water surface, is generated by a few hundred lines of Three.js and GLSL. Nothing is downloaded except the source. Drag to orbit, or switch scenes.

// code → 3D
Read the program, see the world it builds.
One case per language across all four tracks. The left pane is the actual source file that produced the render on the right; the GLSL shader is compiled and running in your browser, and the URDF assets move between their real joint states.

NACRE · sectioned nautilus
57 chambers, nacreous layers and fine growth lines, on a stand that actually touches the shell.
// 3DCodeVerse-2M
A large-scale corpus of executable 3D code.
Programs come from public datasets, from converting specialized representations (DeepCAD sequences and Articraft geometry become standalone CadQuery), and from running the harness itself. Each program ships with the sources and assets it needs to execute, renders, three kinds of captions (instruction, visual description, construction procedure) and its verification record.

Packaged catalogue
269,597 programs across ten collections, expanded into 2,013,629 render and caption entries: the 2.03M code–condition pairs. 3DCodeBench factories, instances and test splits are held out for evaluation.
// 3DCodeVerse Harness
A data engine that checks its own work.
Visually plausible is not the same as geometrically valid. The harness pairs a VLM judge with deterministic measurements, turns every finding into a targeted edit, and loops without human feedback, scaling at inference time to projects with thousands of lines of code.

- 01Plan
An LLM planner turns the prompt into components, dimensions, materials, spatial relations and acceptance criteria.
- 02Build
The plan becomes a DAG of coding and assembly tasks. Sub-agents own separate files and write them in parallel.
- 03Check
Execution logs, then spatial tools: contacts, collisions, joint sweeps and placement, each reported per part in millimetres.
- 04Judge
A VLM judge reads labelled multi-view renders against the plan. Measurements cap the score and veto claims they rule out.
- 05Fix
Findings become localized repair tasks for a fix agent. Every round is kept, and the highest-scoring program wins.
auto-generated feedback → targeted edit
“Belly plates sink 3.2 mm into the clouds (≤ 2 mm).”
“Scales read as domed buttons; flatten the dome to 0.05w.”
adversarial loop engineering
During development the judge and the generator are refined in turns. The judge is tuned on a fixed bank of stored artifacts to reduce score variance and contradictions with measurements; then, with the judge frozen, harness designs are compared on matched prompts, trading judged quality against workflow failures.
// results
Same model, better harness. Same model, better data.
Across four backbones, outputs from each model's native harness are rarely preferred, while the same model inside the 3DCodeVerse harness wins most comparisons. Fine-tuning open Qwen models on 3DCodeVerse-2M turns mostly failing programs into runnable ones.
Human preference, 3D objects
Share of blind side-by-side comparisons won; ties split equally
Judge score, 3D objects
Gemini 3.1 Pro rubric score, 0–100; failed generations score 0
Execution rate after fine-tuning on 3DCodeVerse-2M
% of programs that run in the dialect's pinned runtime; T = 0.7, one seed
Show as a table
| Model | 3DCodeBench(212) | Blender(103) | CadQuery(200) | OpenSCAD(50) | GLSL(200) | Three.js(40) |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | 0.0 | 1.9 | 10.5 | 50.0 | 33.0 | 72.5 |
| + SFT | 84.9 +84.9 | 90.3 +88.4 | 98.5 +88.0 | 86.0 +36.0 | 56.0 +23.0 | 74.5 +2.0 |
| Qwen3.8-27B | 51.9 | 74.8 | 65.5 | 80.0 | 57.0 | 72.5 |
| + SFT + DPO | 92.9 +41.0 | 93.2 +18.4 | 98.5 +33.0 | 90.0 +10.0 | 71.5 +14.5 | 92.5 +20.0 |
The preference and judge charts cover 100 text-to-3D object prompts. Human preference also reaches 78–88% on articulated objects, 50 scene prompts and 50 graphics prompts.


Introducing 3DCodeVerse: a 45-second tour of ten code-authored worlds. Sound on.
// gallery
24 worlds, each one a program.
Code-authored showcase scenes. Geometry, procedural materials, lighting, camera and animation are all kept as executable source. Hover to play a clip, and click to read the code next to the render.
These are showcase studies written with the harness and polished over several rounds. They are not the random samples or model comparisons behind the reported results.
// citation
BibTeX
@article{gao2026_3dcodeverse,
title = {3DCodeVerse: Exploring 3D Worlds as Code},
author = {Gao, Yipeng and Zeng, Ziyao and Huang, Jinfa and Li, Yijiang and Xiong, Haolin and Qin, Wang and Yu, Xinlei and Yao, Youheng and Zhang, Jie and Lin, Kun-Yu and Liu, Rong and Zhao, Yinyuan and Zhao, Yajie and Guo, Meiqi and Luo, Jiebo and Shu, Lei and Ge, Yunhao and Hu, Xinyu and Itti, Laurent},
journal = {arXiv preprint},
year = {2026}
}Related: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling via Code.