Yaofeng Su

Yaofeng Su

B.S. in Computer Science · Fudan University

About

I am an undergraduate at Fudan University (GPA 3.84/4.00). Currently, I am an intern at Joy Future Academy, JD, researching streaming audio-visual generation, DMD distillation algorithms, and world models with world model memory. Our team's JoyAI-Echo achieves minute-level coherent audio-video story generation, and my OmniForcing (ECCV 2026, cited by the Alibaba Wan Team) was presented as an invited talk at Lightricks (LTX-Video Team) to their CTO and Director of GenAI Research. In 2025, I was an exchange student at UNC Chapel Hill, supervised by Prof. Huaxiu Yao, where I contributed to SimpleMem (ICML 2026) based on deep analysis of long-term conversation characteristics and memory structures.

I am seeking PhD opportunities in multimodal foundation models for Fall 2027. I aspire to conduct research grounded in solid theoretical insights, with a particular interest in physically faithful world representations and enabling real-world tasks. If you have any relevant openings or would like to discuss potential collaboration, please feel free to contact me at yaofengsu@outlook.comI would truly appreciate it.

Research Interests

My research interest lies in multimodal generative models as a means to build faithful representations of the world. Through reinforcement learning, distillation, and agent frameworks, I have contributed to applications spanning audio-visual generation, autonomous driving, and video understanding. I am particularly drawn to developing principled approaches that bridge generative modeling with physically grounded world understanding.

Multimodal Generative Models Reinforcement Learning Diffusion Models & SDE LLM/VLM Distillation AI Agent World Models Physically Faithful Representations Embodied Intelligence

Education

B.S. in Computer Science, GPA 3.84/4.00· Fudan University
2023 – 2027
Exchange Student, Dean's List· UNC Chapel Hill
2025

Invited Talks

Invited Talk, Lightricks (LTX-Video Team)
May 2026
Invited by Lightricks to present OmniForcing to the core LTX-Video research team, including the company's CTO and the Director of GenAI Research. Discussed the design and distillation of real-time streaming joint audio-visual generation built on top of the LTX-2 foundation model, covering architectural innovations and ongoing collaboration on LTX-2.3 adaptation.

Publications

OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
Yaofeng Su*, Yuming Li*, Zeyue Xue, Jie Huang, Siming Fu, Haoran Li, Haoyang Huang, Nan Duan
ECCV 2026, Oral
JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
Echo Team @ Joy Future Academy, JD
Tech Report, 2026
SimpleMem: Efficient Lifelong Memory for LLM Agents
Jiaqi Liu*, Yaofeng Su*, Peng Xia, Siwei Han, Zeyu Zheng, Cihang Xie, Mingyu Ding, Huaxiu Yao
ICML 2026
Learning a Unified Risk Map for Autonomous Driving in Partially Observable Environments
Jie Jia*, Yaofeng Su*, Zeyu Bao, Yun Hong, Bingzhao Gao, Zhongxue Gan, Wenchao Ding
IEEE RA-L 2026
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
Yiyang Zhou*, Yangfan He*, Yaofeng Su, Siwei Han, Joel Jang, Gedas Bertasius, Mohit Bansal, Huaxiu Yao
NeurIPS 2025

CV

Download CV (PDF)