About
I am an undergraduate at Fudan University (GPA 3.84/4.00).
Currently, I am an intern at Joy Future Academy, JD, researching streaming audio-visual generation,
DMD distillation algorithms, and world models with world model memory.
Our team's JoyAI-Echo achieves minute-level coherent audio-video story generation,
and my OmniForcing (ECCV 2026, cited by the Alibaba Wan Team) was presented as an
invited talk at Lightricks (LTX-Video Team) to their CTO and Director of GenAI Research.
In 2025, I was an exchange student at UNC Chapel Hill, supervised by Prof. Huaxiu Yao,
where I contributed to SimpleMem (ICML 2026)
based on deep analysis of long-term conversation characteristics and memory structures.
I am seeking PhD opportunities in multimodal foundation models for Fall 2027.
I aspire to conduct research grounded in solid theoretical insights, with a particular interest in
physically faithful world representations and enabling real-world tasks.
If you have any relevant openings or would like to discuss potential collaboration,
please feel free to contact me at yaofengsu@outlook.com — I would truly appreciate it.
Research Interests
My research interest lies in multimodal generative models as a means to build faithful representations of the world.
Through reinforcement learning, distillation, and agent frameworks, I have contributed to applications spanning
audio-visual generation, autonomous driving, and video understanding.
I am particularly drawn to developing principled approaches that bridge generative modeling with physically grounded world understanding.
Multimodal Generative Models
Reinforcement Learning
Diffusion Models & SDE
LLM/VLM Distillation
AI Agent
World Models
Physically Faithful Representations
Embodied Intelligence
Education
B.S. in Computer Science, GPA 3.84/4.00· Fudan University
2023 – 2027
Exchange Student, Dean's List· UNC Chapel Hill
2025
Invited Talks
Invited Talk, Lightricks (LTX-Video Team)
May 2026
Invited by Lightricks to present OmniForcing to the core LTX-Video research team, including the company's CTO and the Director of GenAI Research. Discussed the design and distillation of real-time streaming joint audio-visual generation built on top of the LTX-2 foundation model, covering architectural innovations and ongoing collaboration on LTX-2.3 adaptation.
Publications
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
Yaofeng Su*, Yuming Li*, Zeyue Xue, Jie Huang, Siming Fu, Haoran Li, Haoyang Huang, Nan Duan
ECCV 2026, Oral
JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
Echo Team @ Joy Future Academy, JD
Tech Report, 2026
SimpleMem: Efficient Lifelong Memory for LLM Agents
Jiaqi Liu*, Yaofeng Su*, Peng Xia, Siwei Han, Zeyu Zheng, Cihang Xie, Mingyu Ding, Huaxiu Yao
ICML 2026
Learning a Unified Risk Map for Autonomous Driving in Partially Observable Environments
Jie Jia*, Yaofeng Su*, Zeyu Bao, Yun Hong, Bingzhao Gao, Zhongxue Gan, Wenchao Ding
IEEE RA-L 2026
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
Yiyang Zhou*, Yangfan He*, Yaofeng Su, Siwei Han, Joel Jang, Gedas Bertasius, Mohit Bansal, Huaxiu Yao
NeurIPS 2025