3DCodeVerse

Exploring 3D Worlds as Code

Yipeng Gao1,*,†Ziyao Zeng2,*Jinfa Huang3,*Yijiang Li4,*

Haolin Xiong1Wang Qin1Xinlei Yu1Youheng Yao2Jie Zhang5Kun-Yu Lin6Rong Liu1Yinyuan Zhao7

Yajie Zhao1Meiqi Guo8Jiebo Luo3Lei Shu8Yunhao Ge1Xinyu Hu9Laurent Itti1

1University of Southern California2Yale University3University of Rochester4University of California, San Diego5Columbia University6Nanyang Technological University7University of California, Santa Cruz8Google DeepMind9Philo Labs

* Equal contribution† Project Lead

PaperarXiv soonCodeDatasetVideo

The Jade Guardian, running live. Every mesh, material and light, and the water surface, is generated by a few hundred lines of Three.js and GLSL. Nothing is downloaded except the source. Drag to orbit, or switch scenes.

2.03M
code–condition pairs
from 269,597 executable programs
7
languages & runtimes
bpy · CadQuery · OpenSCAD · Three.js · URDF · GLSL · OpenGL
4
tracks
objects · articulated · scenes · graphics
90%
human preference
vs. Claude Design, same Claude Opus 5.5 backbone
92.9%
execution rate
fine-tuned 27B on 3DCodeBench, up from 51.9%
3DCodeVerse showcase: scenes, graphics, articulated objects and 3D objects, all built by code

// code → 3D

Read the program, see the world it builds.

One case per language across all four tracks. The left pane is the actual source file that produced the render on the right; the GLSL shader is compiled and running in your browser, and the URDF assets move between their real joint states.

nacre.py· 174 linesBlender Python
output· rendered by the runtime
NACRE · sectioned nautilus

NACRE · sectioned nautilus

57 chambers, nacreous layers and fine growth lines, on a stand that actually touches the shell.

// 3DCodeVerse-2M

A large-scale corpus of executable 3D code.

Programs come from public datasets, from converting specialized representations (DeepCAD sequences and Articraft geometry become standalone CadQuery), and from running the harness itself. Each program ships with the sources and assets it needs to execute, renders, three kinds of captions (instruction, visual description, construction procedure) and its verification record.

Curation and composition of 3DCodeVerse-2M
Left: programs collected from public datasets and generated with the harness. Middle: a coding agent with modeling skills and inspection tools (3D IoU, collision, overlap, distance) builds, judges and refines each program. Right: the corpus spans objects, articulated objects, scenes and graphics.

Packaged catalogue

269,597 programs across ten collections, expanded into 2,013,629 render and caption entries: the 2.03M code–condition pairs. 3DCodeBench factories, instances and test splits are held out for evaluation.

ilabai/3dcodeverse on Hugging Face
DeepCAD · DeepCAD → CadQuery
128,201
Shadertoy · GLSL
119,961
Articraft · CadQuery
6,882
Bioinspired3D · Blender Python
6,343
Articraft · URDF
6,146
Thingiverse · OpenSCAD
1,193
3DCodeBench factories · Blender Python
486
Agent-written · Three.js
280
Agent-written · Blender Python
80
Three.js repositories · Three.js / R3F
25

// 3DCodeVerse Harness

A data engine that checks its own work.

Visually plausible is not the same as geometrically valid. The harness pairs a VLM judge with deterministic measurements, turns every finding into a targeted edit, and loops without human feedback, scaling at inference time to projects with thousands of lines of code.

Overview of the 3DCodeVerse Harness
An LLM planner defines the target and a DAG coordinates sub-agents to code and assemble the program (left). Execution checks, spatial measurements and rubric-guided VLM judgments feed a fix agent (center, bottom). Successive rounds are scored and the best program is kept (right).
  1. 01
    Plan

    An LLM planner turns the prompt into components, dimensions, materials, spatial relations and acceptance criteria.

  2. 02
    Build

    The plan becomes a DAG of coding and assembly tasks. Sub-agents own separate files and write them in parallel.

  3. 03
    Check

    Execution logs, then spatial tools: contacts, collisions, joint sweeps and placement, each reported per part in millimetres.

  4. 04
    Judge

    A VLM judge reads labelled multi-view renders against the plan. Measurements cap the score and veto claims they rule out.

  5. 05
    Fix

    Findings become localized repair tasks for a fix agent. Every round is kept, and the highest-scoring program wins.

auto-generated feedback → targeted edit

“Belly plates sink 3.2 mm into the clouds (≤ 2 mm).”
“Scales read as domed buttons; flatten the dome to 0.05w.”

adversarial loop engineering

During development the judge and the generator are refined in turns. The judge is tuned on a fixed bank of stored artifacts to reduce score variance and contradictions with measurements; then, with the judge frozen, harness designs are compared on matched prompts, trading judged quality against workflow failures.

// results

Same model, better harness. Same model, better data.

Across four backbones, outputs from each model's native harness are rarely preferred, while the same model inside the 3DCodeVerse harness wins most comparisons. Fine-tuning open Qwen models on 3DCodeVerse-2M turns mostly failing programs into runnable ones.

Human preference, 3D objects

Share of blind side-by-side comparisons won; ties split equally

Native harness3DCodeVerse harness
Gemini 3.8 Flash
vs. Gemini CLI
13%
87%
Claude Opus 5.5
vs. Claude Code
21%
79%
GPT-6 Astra
vs. Codex
31%
69%
GPT-6 Luna
vs. Codex
11%
89%
Claude Opus 5.5
vs. Claude Design
10%
90%

Judge score, 3D objects

Gemini 3.1 Pro rubric score, 0–100; failed generations score 0

Native harness3DCodeVerse harness
Gemini 3.8 Flash
vs. Gemini CLI
86.6
Claude Opus 5.5
vs. Claude Code
88.5
GPT-6 Astra
vs. Codex
87.7
GPT-6 Luna
vs. Codex
85.9
Claude Opus 5.5
vs. Claude Design
87.7
0255075100

Execution rate after fine-tuning on 3DCodeVerse-2M

% of programs that run in the dialect's pinned runtime; T = 0.7, one seed

Qwen3.8-27B (pretrained)Qwen3.8-27B + SFT + DPO
3DCodeBench
212 tasks
92.9%
Blender
103 tasks
93.2%
CadQuery
200 tasks
98.5%
OpenSCAD
50 tasks
90%
GLSL
200 tasks
71.5%
Three.js
40 tasks
92.5%
0255075100
Show as a table
Model3DCodeBench(212)Blender(103)CadQuery(200)OpenSCAD(50)GLSL(200)Three.js(40)
Qwen3.5-9B0.01.910.550.033.072.5
  + SFT84.9 +84.990.3 +88.498.5 +88.086.0 +36.056.0 +23.074.5 +2.0
Qwen3.8-27B51.974.865.580.057.072.5
  + SFT + DPO92.9 +41.093.2 +18.498.5 +33.090.0 +10.071.5 +14.592.5 +20.0

The preference and judge charts cover 100 text-to-3D object prompts. Human preference also reaches 78–88% on articulated objects, 50 scene prompts and 50 graphics prompts.

Qualitative comparisons and automated refinement
(a) The same prompts and backbones with the native harness (Without) and with the 3DCodeVerse harness (With). (b) Selected rounds of automated refinement; numbers are rubric scores on [0, 1].
Applications of 3D code in robotics and video generation
Applications. Code-built articulated assets and scenes go to robot simulation (MuJoCo) and act as spatial and temporal guides for video generation models.

Introducing 3DCodeVerse: a 45-second tour of ten code-authored worlds. Sound on.

// citation

BibTeX

@article{gao2026_3dcodeverse,
  title   = {3DCodeVerse: Exploring 3D Worlds as Code},
  author  = {Gao, Yipeng and Zeng, Ziyao and Huang, Jinfa and Li, Yijiang and Xiong, Haolin and Qin, Wang and Yu, Xinlei and Yao, Youheng and Zhang, Jie and Lin, Kun-Yu and Liu, Rong and Zhao, Yinyuan and Zhao, Yajie and Guo, Meiqi and Luo, Jiebo and Shu, Lei and Ge, Yunhao and Hu, Xinyu and Itti, Laurent},
  journal = {arXiv preprint},
  year    = {2026}
}

Related: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling via Code.