Generative World Models
Generative World Model for 7V Closed-Loop Driving Simulation
A real-time streaming 7-camera world model powered by causal DiT architecture, 3-stage training paradigm, and interactive vector-to-neural closed-loop simulation.
Overview
What is this project about?
Real-world test drives cannot efficiently scale to rare long-tail hazards, while classical graphics simulators suffer from severe sim-to-real domain gaps. Existing surround video diffusion models are non-causal and require 30+ iterative denoising steps, making real-time closed-loop policy evaluation impossible.
A real-time streaming 7V surround world model with block-causal spatio-temporal attention, cross-view epipolar conditioning, and a 3-stage training pipeline (Bidirectional Teacher → Causal Student → DMD Self-Forcing Distillation) that generates photorealistic 7-camera video streams from interactive BEV vector layouts.
Achieved continuous streaming rollout with only 4 denoising steps per chunk, supporting interactive closed-loop simulation across highway merging, dense night traffic, cutting-in maneuvers, and controllable weather synthesis (dusk, rain, storm, snow).
My role. Designed the 7V streaming DiT network architecture, formulated the cross-view causal attention mechanism, and developed the end-to-end 3-stage distillation and closed-loop vector-to-video rollout platform.
01 · System Architecture & Methodology
Streaming 7V Generative Simulation Architecture
A complete closed-loop neural simulation framework translating high-level policy actions into photorealistic, spatially and temporally coherent 7-camera surround video streams.
Streaming World Model DiT Block
28-layer DiT block with AdaLN modulation, intra-view temporal attention with KV-caching, localized cross-view spatial attention, and cross-attention text/vector conditioning.
Three-Stage Progressive Training
Stage 1 (L1b): Bidirectional Teacher pretraining.
Stage 2 (L2a): Causal Student conversion with block-causal mask.
Stage 3 (L3): Self-Forcing DMD distillation (35 steps → 4 steps).
02 · Model Evolution & Data Engine
Teacher Denoising & Causal Student Conversion
Comparing full-window bidirectional teacher generation against streaming causal student rollout and rare-case safety scenario synthesis.
Surround Generation Walkthrough — Case 1
Surround Generation Walkthrough — Case 2
Rare Interactive Safety Case — Example 1
Rare Interactive Safety Case — Example 2
Causal Block Rollout — Example 1
Causal Block Rollout — Example 2
03 · Stage 3 Self-Forcing & Interactive Rollout
Streaming Closed-Loop Traffic Scenarios
Final self-forcing distilled model executing continuous real-time neural rollouts across interactive driving maneuvers in complex traffic environments.
Daytime Highway Overtaking & Lane Change
Nighttime Congested Highway Traffic
Highway Confluence & Ramp Merging
Dusk Generation vs Ground Truth Benchmark
04 · Zero-Shot Weather Transfer
Multi-Weather Controllable Generation
Demonstrating zero-shot style transfer on the identical trajectory layout across harsh weather conditions to test perception robustness.