A generative simulation platform built on Cosmos-Transfer2.5: a 7-camera surround world model, real-map scenario generation, and a 4-step distilled sampler that makes surround rollouts fast enough to sit inside a closed loop.
A two-level driving simulator: a vector world model that decides what the traffic does, and a sensor-level pipeline that decides what the cameras see — reconstruction, generated traffic, and a mask-guided video editor in one closed loop.
A one-stage, pure-vision end-to-end driving POC: eight surround cameras lifted into a single BEV feature, three perception heads and a Diffusion-Flow planner trained together, with no hand-designed 3D interface in between.
Production perception on a mid-trim J6E / J6M platform: a static OneModel driving every static element from one shared BEV feature, a 4D-sparse model that fuses detection and tracking, and a latency-compression effort that took inference from ~42.65 ms to ~13.88 ms.
An end-to-end driving system fusing 11 surround cameras and LiDAR under a sparse-centric paradigm. I owned two pieces: a fused CUDA aggregation operator, and the AI planner that emits motion and planning in parallel.
A controllable surround-view driving generator: 3D boxes and maps become spatial conditions, text / reference frames / lanes / calibration become condition tokens, and a UNet diffusion backbone turns them into cross-camera-consistent 4V / 7V / 11V video.
Hozon Auto × SJTU IRMV · PhiGent Robotics · Perception Team Leader · 3D Perception Algorithm Engineer
Two eras of autonomous-driving auto-labeling: a vision-only 4D pipeline built with Hozon Auto and SJTU IRMV, then a multi-modal 4D auto-labeler and a production pure-LiDAR 3D detector at PhiGent Robotics.
PhiGent Robotics · 3D Scene Flow Algorithm Engineer
An unsupervised 3D scene-flow auto-labeler that assigns a motion vector to every LiDAR point and occupancy cell, distilled into an ultra-light production head and deployed through one ONNX graph to both TensorRT (Orin) and Horizon J6E.
Road-surface perception for a road-preview ('magic-carpet') suspension feature: segment manhole covers and speed bumps reliably in the wild, then compress and quantize the model to INT8 for TDA4 edge inference.
A short introduction film I made for my 2023 master's graduation — a compact tour of my research focus, the labs and mentors I worked with, and the perception and 3D/4D systems I built along the way.
The safety-critical perception stack for a robot that has to drive itself off a transport vehicle, reach the lawn, mow it, and come back — four modules from ramp detection to an MCU-deployed BEV safety net.
The Future Laboratory of the Second Aerospace Academy · Perception and Simulation Developer
A single end-to-end network that fuses RGB, LiDAR and infrared through attention, then jointly solves geometry–semantic mapping, unsupervised depth and odometry, detection and tracking, and closed-loop behaviour decisions.
researche2eperception
Esc
Type to search across publications, projects, and blog notes.