About Me

I am a PhD student in Computer Science and Engineering at The Chinese University of Hong Kong, advised by Prof. Yu Cheng. My research focuses on multimodal foundation models, particularly unified vision-language models and world/action models.

Before joining CUHK, I received my M.S. in Computer Science from the National University of Singapore and my B.S. in Computer Science from Shanghai Jiao Tong University, where I was advised by Prof. Junchi Yan.

Selected Research

Matrix Game 3.5

A long-horizon interactive world model with persistent memory. I led the design and implementation of its memory module, including dynamic object filtering and object-token conditioning for scene and protagonist consistency.

Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models

BMVC 2026. We systematically study how visual understanding and generation benefit each other in unified VLMs, and identify input-output visual-space alignment as a key factor in cross-task knowledge transfer.

Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models

AAAI 2026. We compress visual representation sequences by 4x, achieving a 10x inference speedup while reducing memory consumption and improving autoregressive video generation quality.

CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling

EMNLP 2025. We upcycle complementary CLIP models into a mixture-of-experts vision encoder with minimal computational overhead, improving performance across downstream MLLM benchmarks.

News

  • 2026: Our work on cross-task generalization between understanding and generation in unified vision-language models was accepted to BMVC 2026.

  • 2026: Matrix Game 3.5 was released.