I work at the intersection of self-evolving agents, multimodal foundation models, and visual content generation — connecting research ideas to open-source systems and to products people actually use.
At MSRA I lead the recursive self-improvement (RSI) research direction, spanning agent harnesses, tasks and rewards, data pipelines, and model training, and extending toward AI4AI: using AI systems to improve AI research and development itself. I am also a long-term contributor to Microsoft's first-party multimodal models, including Phi-3.5 Vision, Phi-4-mini, and the next-generation MAI model series.
40+ peer-reviewed papers (20+ as first or corresponding author) at CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, and AAAI · 10+ international patents as primary inventor · Technology transferred into GitHub Copilot, Azure AI Foundry, and Office.
- SkillOpt — Trains reusable natural-language skills for frozen LLM agents through trajectory-driven optimization and validation-gated updates. Best or tied-best across all 52 evaluated settings.
- Resource2Skill — Distills tutorials, videos, code, and other human-created multimodal resources into reusable, executable agent skills.
- SkillLens — A systematic study of how model-generated skills are produced, consumed, and reused across the skill lifecycle.
- LLM2CLIP — Uses large language models as textual teachers to improve vision-language representation learning. AAAI 2026 Outstanding Paper Award and WAIC 2026 Youth Outstanding Paper Nomination Award; 300K+ Hugging Face downloads, and the technique now serves as the visual-pretraining approach for Phi-4-mini.
- World-R1 — Reinforces 3D constraints in text-to-video generation with camera-aware initialization and 3D-aware rewards. ICML 2026.
- Latent Spatial Memory — Persistent latent 3D scene memory for video world models: 10.57× faster generation and 55× lower 3D-cache memory.
- RAS — Region-adaptive, training-free sampling for efficient diffusion-transformer inference. CVPR 2026.
- BizGenEval — A systematic benchmark for commercial visual content generation, now used as a reward and evaluation system in Microsoft PowerPoint.
- AVGen-Bench — Task-driven, multi-granular evaluation for text-to-audio-video generation.
- Aug 2026 · Paper acceptances: 2 papers accepted to EMNLP 2026.
- Jul 2026 · Paper acceptances: 5 papers accepted to ECCV 2026.
- Jul 2026 · Invited talk, "From Skill Understanding to Skill Optimization: Controllable Text-Space Training for Self-Evolving Agents": presented at CCF SPP and at Zhiyuan College, Shanghai Jiao Tong University.
- Jun 2026 · Microsoft Research Asia Live Talk: An In-Depth Look at SkillOpt.
- Jun 2026 · China Agent Conference: gave an invited talk and chaired the technical workshop on Agent Skills.
- Jun 2026 · Galaxy Workshop: invited talk, "How Do Agent Skills Self-Evolve?".
- Jun 2026 · Paper acceptances: 6 papers accepted to ICML 2026.
- Jan 2026 · LLM2CLIP: received the AAAI 2026 Outstanding Paper Award (media coverage), and presented in a BAAI Community live talk on the frontiers of multimodal representation.
- Area Chair: NeurIPS, ICML, and ICLR
- Senior Program Committee Member: AAAI
- Program Committee / Reviewer: CVPR, ICCV, ECCV, SIGGRAPH, WAICA, BMVC, NeurIPS, ICLR, ICML, ACL, EMNLP, AISTATS, and WACV
- Journal Reviewer: IJCV, IEEE TMM, and TMLR
I maintain long-standing research collaborations with Tsinghua, Peking University, Fudan, Shanghai Jiao Tong, and Tongji, and have mentored 30+ master's and Ph.D. students — many now pursuing Ph.D.s or faculty positions at top North American universities, or working at NVIDIA, Meta, Alibaba Qwen, and Tencent Hunyuan. Always open to conversations about self-evolving agents and multimodal generation — reach me at yifyang29@gmail.com.


