My current research focuses on Multimodal Generation and Understanding, with a particular interest in language, video, image generation and multimodal agent memory. Previously, I explored Optimal Transport theory for Unsupervised Deep Clustering and 3D Semi-Supervised Learning, with a focus on imbalanced and open-world scenarios.
Generally speaking, my research aim is advancing intelligent systems that can simulate, understand, and interact with the world while efficiently leveraging prior knowledge to generalize across diverse real-world scenarios. - How far are we from the genuine advent of Artificial General Intelligence?
I am open to collaboration and discussion. If you are interested, don't hesitate to reach out!
Publications
Note: * indicates equal contribution.
Project on Long Video Understanding 【Life-long Visual Memory】
How to achieve infinite multimodal agent memory? Ongoing
Project on Diffusion Language Models 【Pretrain】
Is autoregression the final solution for LLM? Under Review