I'm a fourth-year Ph.D. student in Computer Science at Shanghai Jiao Tong University (SJTU),
advised by Prof. Lizhuang Ma in the Digital Media & Computer Vision Laboratory (DMCV).
I also receive supervision from Dr. Xin Tan, who is based at East China Normal University (ECNU).
Prior to starting my Ph.D., I received my Bachelor's degree in Computer Science from Beihang University (BUAA).
I also worked as an intern at Baidu.
My research interests involve multimodal large language models (MLLM) and 3D vision.
I am currently interested in MLLM post-training, agentic RL, and world models.
I have previously conducted some work in 2D/3D scene understanding, especially in scene parsing and 3D Gaussian splatting.
2026.08: I am currently looking for collaboration / internship opportunities related to MLLM. Feel free to contact me!
A feed-forward network that reconstructs language-embedded 3D Gaussians from arbitrary uncalibrated and unposed images with a compact semantic representation that avoids per-Gaussian language embedding and significantly reduces storage overhead.
A sparse-to-dense lifting pipeline that bridges point clouds and 3D Gaussian Splatting via a one-step diffusion model, enabling high-quality 3DGS reconstruction with minimal inputs.
A novel setting called Generalized Category Discovery in Semantic Segmentation (GCDSS). Given prior knowledge from a labeled
set of base classes, our method aims to segment unlabeled images that contain pixels of the base class or novel class.