A unified multimodal world model

Puffin-World

Scaling a Unified Multimodal Model
with Native 3D World States

1S-Lab, Nanyang Technological University 2University of Michigan 3Beijing Jiaotong University 4ACE Robotics

Preprint · 2026

Puffin-World physical-world perception and free-viewpoint spatial simulation Puffin-World long-trajectory, extreme-rotation and text-driven world generation Puffin-World native appearance, geometry and physics states with 3D reconstruction

Beyond pixels, toward worlds

One model that can perceive, simulate, and build the 3D world

Puffin-World represents the physical world through three complementary native states—physics, geometry, and appearance. A single unified model connects physical-world perception, free-viewpoint spatial simulation, and 3D world modeling.

Overview of Puffin-World modalities, tasks, and unified understanding, generation, and reconstruction capabilities
Puffin-World unifies camera-physics understanding, free-viewpoint simulation, and 3D world generation and reconstruction through native appearance, geometry, and physics states.

01 · Physics

Physics World State

Gravity-aware camera understanding and physically consistent trajectory propagation across challenging motions.

02 · Geometry

Geometry World State

Dense spatial structure for underlying geometry and native 3D reconstruction.

03 · Appearance

Appearance World State

High-fidelity visual content that remains coherent as the camera moves through a generated world.

Interactive 3D world explorer

Move through the world, not just around an image

Switch between the generated worlds to inspect appearance, gravity, latitude, and geometry together with the corresponding interactive 3D reconstruction.

Appearance world stateGenerated RGB trajectory
Physics world stateGravity / up direction
Physics world stateLatitude / horizon alignment
Geometry world stateDense depth trajectory
3D ReconstructionInteractive colored point cloud
670K points
Loading 3D reconstruction
Interactive 3D reconstruction

Drag orbit · Scroll zoom · Right-drag move

The reconstruction gently sweeps left and right while idle; interact at any time to inspect the world freely.

Diverse Actions: Long Trajectory, Extreme Rotation, and Compound Motions
Four-state generated world sequence 01
Four-state generated world sequence 02
Four-state generated world sequence 03
Four-state generated world sequence 04
Four-state generated world sequence 05
Four-state generated world sequence 06
Four-state generated world sequence 07
Four-state generated world sequence 08
Four-state generated world sequence 09
Compound-motion generated world sequence 01
Three-state generated world sequence 01
Extreme-rotation generated world sequence 06
Three-state generated world sequence 04
Extreme-rotation generated world sequence 09
Three-state generated world sequence 07
Compound-motion generated world sequence 04
Compound-motion generated world sequence 02
Three-state generated world sequence 02
Extreme-rotation generated world sequence 07
Three-state generated world sequence 05
Extreme-rotation generated world sequence 10
Three-state generated world sequence 08
Compound-motion generated world sequence 03
Three-state generated world sequence 03
Compound-motion generated world sequence 05
Three-state generated world sequence 06
Extreme-rotation generated world sequence 08
Extreme-rotation generated world sequence 11

For each sample shown above, only the initial appearance world state (the first RGB view) is provided as input; all remaining world states are generated by Puffin-World.

Unified architecture

Native world states, learned in one framework

Puffin-World combines a geometry-aligned vision encoder, an LLM, a diffusion model, and a lightweight connector. It supports understanding, generation, and reconstruction without relying on task-specific external geometry modules.

Puffin-World network architecture
The network mainly comprises 3D world understanding, generation, and reconstruction, formulating native 3D world states within one framework.
9-channel condition

Omni-Camera Representation

A gravity-aware absolute perspective field is paired with ray-based relative geometry, enabling precise control over camera intrinsics, orientation, and trajectory.

Physically grounding consistency

Physics Propagation

The model propagates physical constraints through generated trajectories, improving visual consistency and preserving gravity under challenging rotations and long camera motion.

Unified training

Understanding + Generation

Causal reasoning tokens and asymmetric diffusion attention let the same model interpret camera geometry and synthesize world-consistent observations.

Four capability families

A complete loop from perception to interaction

01

Physical-World Perception

Understands camera parameters, spatial relationships, and the physical principles behind an observation.

02

Free-Viewpoint Simulation

Generates observations from explicit camera poses across roll, pitch, and intrinsic parameters.

03

3D World Modeling

Supports image-to-3D, text-to-3D, long rotations, native states, and direct 3D reconstruction.

04

Closed-Loop Applications

Enables interleaved conversation, world exploration, and self-calibrated interaction with generated environments.

Camera-to-World Understanding Results

Strong camera-to-world perception from a single image

Puffin-World estimates absolute camera roll, pitch, and vertical field-of-view (vFoV), grounding visual observations in a gravity-aligned camera-to-world frame. Across Stanford2D3D, MegaDepth, TartanAir, and LaMAR, it leads every median-error metric and most AUC metrics.

12/12Best median-error results
33/36Best AUC metrics, including ties
0.26°Lowest roll error · LaMAR
1.62°Lowest vFoV error · Stanford2D3D

Evaluation results on camera-to-world understanding

Median error and AUC at 1°, 5°, and 10° thresholds across four public benchmarks.

DatasetApproachRoll [degrees]Pitch [degrees]vFoV [degrees]
Error ↓AUC ↑Error ↓AUC ↑Error ↓AUC ↑
10°10°10°
Stanford2D3DDeepCalib1.5933.863.979.22.5821.646.965.76.678.120.637.6
Perceptual2.0826.853.870.73.1721.541.857.813.842.87.716.1
CTRL-C3.0423.243.056.93.4318.338.653.88.507.718.231.5
MSCC3.4313.536.857.32.6422.645.060.55.819.623.841.6
ParamNet1.1444.673.984.81.9429.256.773.19.015.814.327.8
SVA-21.724.625.8-15.419.922.4-6.211.515.2
UVP0.5265.374.679.10.9551.263.069.23.6522.239.551.3
GeoCalib0.4083.191.894.80.9352.374.884.63.2117.440.059.4
AnyCalib†--------2.5521.146.864.6
Puffin-World0.2993.197.498.50.5373.388.894.01.6234.561.176.4
MegaDepthDeepCalib1.4134.665.479.45.1911.927.844.811.145.612.122.9
Perceptual1.0747.972.483.23.4919.839.154.213.402.98.216.8
CTRL-C0.8854.575.084.24.8016.633.246.518.652.05.812.8
MSCC0.9053.172.882.15.7319.033.244.310.806.014.626.2
ParamNet1.1743.470.782.23.9915.434.553.311.013.210.121.3
SVA-31.935.036.2-13.620.624.9-9.416.121.1
UVP0.5169.281.686.94.5921.636.247.410.928.218.729.8
GeoCalib0.3682.690.694.01.9432.453.367.54.4613.631.748.2
AnyCalib†--------3.1419.440.859.1
Puffin0.3284.993.496.21.0847.668.279.42.4223.947.864.1
Puffin-World0.2887.794.296.71.0149.769.379.62.4124.848.865.5
TartanAirDeepCalib1.9524.755.471.53.2716.338.858.58.071.58.827.2
Perceptual2.2423.248.666.72.8623.544.661.515.065.18.917.1
CTRL-C1.6832.859.174.12.3924.648.665.25.6410.725.443.5
MSCC3.5015.037.257.73.4818.838.654.311.184.411.823.0
ParamNet1.6334.559.273.93.0519.442.060.38.216.016.831.6
SVA9.4832.439.644.118.4621.228.834.543.018.816.121.6
UVP0.8952.164.871.92.4836.248.858.69.1515.825.835.7
GeoCalib0.4371.383.889.81.4938.262.976.64.9014.130.447.6
AnyCalib†--------3.6215.536.455.1
Puffin0.4071.786.292.10.9551.068.279.37.4816.328.539.0
Puffin-World0.3180.190.094.20.6760.278.187.22.3426.646.459.5
LaMARDeepCalib1.1544.173.984.84.6810.828.349.810.930.713.024.0
Perceptual1.2940.068.981.62.8321.244.762.617.783.05.310.7
CTRL-C1.2043.570.982.51.9427.654.770.25.649.824.643.2
MSCC1.4439.660.772.83.0220.941.855.714.783.28.316.8
ParamNet0.9351.777.086.02.1527.052.770.214.712.86.814.3
SVA-8.69.29.7-3.45.77.0-1.22.74.1
UVP0.3872.781.885.71.3442.359.969.45.5715.630.643.5
GeoCalib0.2886.492.595.00.8755.076.986.23.0319.141.560.0
AnyCalib†--------2.2524.651.670.5
Puffin0.3880.689.893.50.7161.778.986.43.6217.037.353.1
Puffin-World0.2685.892.595.30.7163.581.288.22.7319.043.059.4

BestSecond best AnyCalib† is specialized for camera-intrinsic estimation and is included as a reference.

Camera-Controllable Generation Results

Precise camera control with strong visual quality

On Puffin-Cam-Bench, Puffin-World combines low angular error with competitive image fidelity, producing visually convincing images that remain closely aligned with the requested camera configuration.

0.84°Median up-vector error
1.26°Median latitude error
0.79°Median gravity error
LowestFID on Puffin-Cam-Bench
Free-viewpoint spatial simulation comparison across image generation models
Free-viewpoint text-to-image generation. The rightmost column shows Puffin-World results; angular values closer to 0° indicate better spatial alignment.

3D World Modeling Results

Build coherent worlds across viewpoints

Puffin-World extends single-view generation to multi-view world modeling with explicit camera trajectories. It supports image-to-3D and text-to-3D generation, large rotations, compound motion, native appearance–geometry–physics prediction, and direct 3D reconstruction.

Qualitative results of Puffin-World 3D world modeling across image-to-3D, text-to-3D, camera control, native world states, and reconstruction
Qualitative results of 3D world modeling. Red boxes indicate model inputs; the remaining views and native world states are generated or predicted by Puffin-World.

Closed-Loop Applications

Perceive, reason, and explore in a closed loop

Puffin-World combines world understanding and generation within one model to support mimic world exploration and self-calibrated world exploration.

Mimic world exploration and self-calibrated world exploration with Puffin-World
Fig. 7. Closed-loop applications of Puffin-World. Mimic world exploration follows a prescribed camera trajectory, while self-calibrated world exploration reasons about the current physical state and predicts corrective camera actions.

Scaling with Puffin-16M

From isolated views to diverse 3D trajectories

Puffin-16M combines dense camera-aware visual supervision with trajectory data that covers challenging pitch, yaw, roll, and compound motion.

15Mvision–language–camera triplets
1Mdiverse camera trajectories
28curated public datasets
44.5Mcamera-labeled images
Overview of Puffin-16M camera-aware images and trajectory data

Large-Scale Camera Grounding

Towards Large-Scale Spatial Intelligence and Physical AI

Puffin-World annotates roll, pitch, and vertical field-of-view for 28 widely used datasets—22 single-image datasets and six sequential or 3D Absolute-Camera datasets—yielding approximately 44.5M camera-grounded images.

Overview of 28 public datasets annotated with camera parameters by Puffin-World
Fig. 8. Overview of 28 public datasets annotated by Puffin-World, covering approximately 44.5M images. All camera-annotated datasets shown here are released in our Hugging Face collection. Open collection ↗

Citation

BibTeX

@article{liao2026puffinworld,
  title   = {Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States},
  author  = {Liao, Kang and Luo, Yihang and Wu, Xiao-Ming and Jin, Linyi and Wu, Size and Lin, Chunyu and Zhao, Yao and Wang, Fei and Li, Wei and Loy, Chen Change},
  journal = {Preprint},
  year    = {2026}
}