🎯 I am seeking Ph.D. opportunities for Fall 2027 in 3D/4D scene representation, video generation, and robot world models.
🔬 Research Interests: 3D/4D Scene Representation · Human-Centric Video Generation · Dynamic Visual Perception · Embodied World Models
I build generative and geometric models for interactive humans, dynamic 3D scenes, and robot learning.
👨🎓 I am pursuing an M.Sc. in Electrical Engineering and Information Technology at the Technical University of Munich, after earning an M.Sc. in Computer Science from Tongji University in 2025.
👨🏻💻 For my master's thesis at TUM Visual Computing, I am developing a camera-controlled video diffusion model for controllable synthesis of human motion and scene interactions across changing viewpoints. I am advised by Yu Chi, Jiapeng Tang, and Prof. Matthias Nießner.
🤖 At Agile Robots SE, I work on cross-embodiment video generation that translates egocentric human demonstrations into robot-domain videos for downstream policy learning.
| Work | Venue | Focus |
|---|---|---|
| ViDS · Paper | Preprint, 2026 | 3D face tracking-conditioned video diffusion for expressive, identity-preserving portrait animation; ranked first on 8 of 13 reported VFHQ metrics. |
| CSG-Fusion | ICCV 2025 Workshop E2E3D · 🏆 Best Paper | Matching-based fusion of sparse-view pointmaps into compact, cross-view-consistent 3D Gaussians. |
| LiteTracker · Code | MICCAI 2025 | Accurate online tissue tracking with causal temporal feature reuse; approximately 7× faster than its predecessor and 2× faster than prior state of the art. |
Previously, I worked on 3D scene decomposition and sparse-view reconstruction with TUM-CVG, TUM-DI-LAB, and Oxford VGG; dense video point tracking with ImFusion and TUM CAMP; and large-scale 3D reconstruction at Zhejiang University.
Across these projects, I collaborated with Prof. Daniel Cremers, Prof. Yan Xia, Prof. Chuanxia Zheng, Prof. Nassir Navab, Prof. Benjamin Busam, and Prof. Yiyi Liao.
💬 I am always happy to discuss research ideas, collaboration, and Ph.D. opportunities. Feel free to email me.
From geometry and video to perception–action models of interactive humans and dynamic scenes.


