I am a Ph.D. candidate at the Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University, advised by Prof. Hang Zhao. Before that, I received my Master's degree from the Chinese University of Hong Kong and my Bachelor's degree from Xidian University, where I was advised by Prof. Bo Chen.
My goal is to build machines that perceive, reason about, and act in the physical world. I am currently working on embodied foundation models — large Vision-Language-Action (VLA) models pretrained across embodiments, scenes, and tasks, so that a single model can follow open-ended language instructions and manipulate unseen objects out of the box. My recent work centers on neural action tokenization and efficient autoregressive VLA architectures (FASTer, ActionCodec), which underpin the Galaxea G0/G0.5 model family. Earlier, I worked on end-to-end autonomous driving and online HD map learning (VectorMapNet, Neural Map Prior, StreamMapNet, DriveVLM).