About Me
I am a PhD student in Computer Science and Engineering at The Chinese University of Hong Kong, advised by Prof. Yu Cheng. My research focuses on multimodal foundation models, particularly unified vision-language models and world/action models.
Before joining CUHK, I received my M.S. in Computer Science from the National University of Singapore and my B.S. in Computer Science from Shanghai Jiao Tong University, where I was advised by Prof. Junchi Yan.
Selected Research
Matrix Game 3.5
A long-horizon interactive world model with persistent memory. I led the design and implementation of its memory module, including dynamic object filtering and object-token conditioning for scene and protagonist consistency.
Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models
BMVC 2026. We systematically study how visual understanding and generation benefit each other in unified VLMs, and identify input-output visual-space alignment as a key factor in cross-task knowledge transfer.
Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models
AAAI 2026. We compress visual representation sequences by 4x, achieving a 10x inference speedup while reducing memory consumption and improving autoregressive video generation quality.
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
EMNLP 2025. We upcycle complementary CLIP models into a mixture-of-experts vision encoder with minimal computational overhead, improving performance across downstream MLLM benchmarks.
News
2026: Our work on cross-task generalization between understanding and generation in unified vision-language models was accepted to BMVC 2026.
2026: Matrix Game 3.5 was released.
