About
I am Yifan Yang, a Senior Research SDE at Microsoft Research Asia (MSRA), Shanghai, where I joined in 2021. My research focuses on visual content generation, multimodal foundation models, and general-purpose agentic systems, with a particular emphasis on bridging research innovation, open-source development, and real-world deployment.
I have published over 30 peer-reviewed papers in top-tier venues, including CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, and AAAI, and have filed more than ten international patents. My academic service includes serving as an Area Chair for NeurIPS, ICML, and ICLR, and as a Senior Program Committee member for AAAI.
I have been deeply involved in the development of Microsoft’s Phi model family, including Phi-3 and Phi-4, with several of my techniques transferred into core Microsoft products, including Office and Azure. My work on LLM2CLIP enhances cross-modal representation learning by leveraging large language models; it has been integrated into the Phi-4-mini pretraining pipeline and received the AAAI 2026 Outstanding Paper Award. My recent research also explores self-evolving agents, commercial visual content generation, video world models, text-to-audio-video generation, and structured multimodal reasoning.
Google Scholar (opens in new tab)
|
GitHub (opens in new tab)
If you are interested in internship opportunities or research collaborations, feel free to contact me at yifanyang@microsoft.com.
Selected Open-Source Projects
My open-source work spans self-evolving agents, multimodal representation learning, world models, and efficient visual generation. As of July 2026, the projects highlighted below have collectively attracted more than 17,000 GitHub stars.
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills (opens in new tab) — A text-space optimizer that treats reusable natural-language agent skills as trainable artifacts. It improves a frozen target model through trajectory-driven edits, bounded textual updates, validation gating, and slow/meta optimization, while adding zero model calls at deployment.Open-source impact: As of July 2026, SkillOpt has received 15.1K GitHub stars and 1.4K forks, ranked #3 among Python repositories of the day on Trendshift, shipped as a PyPI package, and was integrated by community projects including gbrain, gbrain-evals, and darwin-skill. It has also been featured by Microsoft Research and covered by VentureBeat, Synced (机器之心), The Decoder, and other technology media.Research result: SkillOpt is best or tied-best in all 52 evaluated model–benchmark–harness settings across six benchmarks, seven target models, and direct-chat, Codex, and Claude Code execution modes.
Paper (opens in new tab) · Project Page (opens in new tab) · Microsoft Research Blog
- LLM2CLIP (opens in new tab) — Uses large language models as powerful textual teachers for CLIP, strengthening visual representations and text–image alignment for long, dense, and multilingual descriptions. The repository releases training code and model checkpoints. The technique was integrated into the Phi-4-mini pretraining pipeline and received the AAAI 2026 Outstanding Paper Award.Paper (opens in new tab) · Project Page (opens in new tab) · Model Zoo (opens in new tab)
- Resource2Skill (opens in new tab) — Distills human-created multimodal resources—including tutorial videos, articles, code, and reference artifacts—into reusable executable skills that agents can browse, compose, and run in real software environments. Across seven practical authoring domains, it improves the average overall score by +11.9 percentage points over no-skill agents and outperforms strong harness baselines in 26 of 28 main model–domain evaluation cells.Paper (opens in new tab) · Project Page (opens in new tab) · Dataset and Skill Libraries (opens in new tab)
- SkillLens (opens in new tab) — A systematic framework for studying model-generated agent skills across the complete lifecycle from raw experience generation to skill extraction and skill consumption. It provides unified pipelines and reproducible evaluation across SWE-bench, ALFWorld, SpreadsheetBench, BFCL, and SEAL-0.Paper (opens in new tab) · Project Page (opens in new tab)
- World-R1 (opens in new tab) — Reinforces 3D constraints in text-to-video generation through camera-aware latent initialization, 3D-aware rewards, and periodic decoupled training, improving geometric consistency and camera control without changing the base video model architecture. ICML 2026.Paper (opens in new tab) · Project Page (opens in new tab) · Dataset (opens in new tab)
- Latent Spatial Memory (opens in new tab) — Introduces a persistent 3D scene memory represented directly as latent tokens for video world models, avoiding repeated RGB rendering and re-encoding. The released Mirage system reports 10.57× faster generation and 55× lower 3D-cache memory.Paper (opens in new tab) · Project Page (opens in new tab)
- RAS: Region-Adaptive Sampling for Diffusion Transformers (opens in new tab) — A training-free sampling strategy that dynamically allocates computation to visually complex regions while reusing intermediate results in simpler regions, accelerating diffusion-transformer inference with minimal quality loss. CVPR 2026.Paper (opens in new tab) · Project Page (opens in new tab)
Publications
* denotes co-first author; † denotes corresponding author. Publication titles link to arXiv where available. Three works without a verified arXiv record link to the official ACM, IEEE, or ACL publication page.
First-author and Corresponding-author Publications
- LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation (opens in new tab)
Weiquan Huang, Aoqi Wu, Yifan Yang†, et al.
AAAI 2026 — Outstanding Paper Award 🏆 - World-R1: Reinforcing 3D Constraints for Text-to-Video Generation (opens in new tab)
Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yang†, et al.
ICML 2026 - AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation (opens in new tab)
Ziwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang†, et al.
ICML 2026 - Video-in-the-loop: Span-grounded Long Video QA with Interleaved Reasoning (opens in new tab)
Chendong Wang, Donglin Bai, Yifan Yang†, et al.
ICML 2026 - A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding (opens in new tab)
Yida Wang, Taiting Lu, Runze Liu, Lanqing Yang, Yifan Yang†, et al.
ICML 2026 - HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models (opens in new tab)
Ziqin Zhou, Yifan Yang†, et al.
AAAI 2026 - VidGuard-R1: AI-generated Video Detection and Explanation via Reasoning Multimodal Language Models and Reinforcement Learning (opens in new tab)
Kyoungjun Park, Yifan Yang†, et al.
ICLR 2026 - Region-adaptive Sampling for Diffusion Transformers (opens in new tab)
Ziming Liu*, Yifan Yang*, et al.
CVPR 2026 - Diffusion²: Turning 3D Environments into Radio Frequency Heatmaps (opens in new tab)
Kyoungjun Park, Yifan Yang†, et al.
CVPR 2026 Findings - Zoomer: Adaptive Image Focus Optimization for Black-box Multimodal Large Language Models (opens in new tab)
Jiaxu Qian, Chendong Wang, Yifan Yang†, et al.
Transactions on Machine Learning Research (TMLR), 2025 - VoLUT: Efficient Volumetric Streaming Enhanced by LUT-based Super-resolution (opens in new tab)
Chendong Wang, Anlan Zhang, Yifan Yang†, et al.
MLSys 2025 - LoRaSC: Expressive and Generalizable Low-rank Adaptation for Large Models via Slow Cascaded Learning (opens in new tab)
Siwei Li, Yifan Yang†*, et al.
EMNLP 2024 Findings - VIGOR: Reviving Cloud Gaming Sessions (opens in new tab)
Zhaoyuan He, Yifan Yang*, et al.
ACM CoNEXT 2024 - Nerve: Real-time Neural Video Recovery and Enhancement on Mobile Devices (opens in new tab)
Zhaoyuan He, Yifan Yang*, et al.
Proceedings of the ACM on Networking (CoNEXT), 2024 - Attentive Mask CLIP (opens in new tab)
Yifan Yang, et al.
ICCV 2023 - ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image Manipulation (opens in new tab)
Yasheng Sun*, Yifan Yang*, et al.
NeurIPS 2023 - Directional Self-supervised Learning for Heavy Image Augmentations (opens in new tab)
Yalong Bai*, Yifan Yang*, et al.
CVPR 2022
Technical Reports and Major Corresponding-author Preprints
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills (opens in new tab)
Yifan Yang†, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, et al.
arXiv preprint arXiv:2605.23904, 2026 — first and corresponding author - OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout (opens in new tab)
Taiting Lu, Kaiyuan Lin, Mingjia Wang, et al., Yifan Yang†, et al.
arXiv preprint arXiv:2607.03261, 2026 - RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources (opens in new tab)
Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang†, et al.
arXiv preprint arXiv:2606.29538, 2026 - Latent Spatial Memory for Video World Models (opens in new tab)
Weijie Wang, Haoyu Zhao, Yifan Yang†, Feng Chen, Zeyu Zhang, et al.
arXiv preprint arXiv:2606.09828, 2026 - From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills (opens in new tab)
Zisu Huang, Jingwen Xu, Yifan Yang†, Ziyang Gong, Qihao Yang, et al.
arXiv preprint arXiv:2605.23899, 2026 - ReasonGen-R1: Chain-of-Thought for Autoregressive Image Generation Models through Supervised Fine-tuning and Reinforcement Learning (opens in new tab)
Yu Zhang, Yunqi Li, Yifan Yang†, et al.
arXiv:2505.24875, under review at ECCV 2026 - MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation (opens in new tab)
Yan Li, Zezi Zeng, Yifan Yang†, et al.
arXiv preprint arXiv:2604.15309, 2026 - BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation (opens in new tab)
Yan Li, Zezi Zeng, Ziwei Zhou, Xin Gao, Muzhao Tian, Yifan Yang†, et al.
arXiv preprint arXiv:2603.25732, 2026 - OmniSch: A Multimodal PCB Schematic Benchmark for Structured Diagram Visual Reasoning (opens in new tab)
Taiting Lu, Kaiyuan Lin, Yuxin Tian, Yubo Wang, Muchuan Wang, Yifan Yang†, et al.
arXiv preprint arXiv:2604.00270, 2026 - Phi-4-mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs (opens in new tab)
Abdelrahman Abouelenin, Atabak Ashfaq, et al., Yifan Yang
arXiv Technical Report, 2025 - Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone (opens in new tab)
Marah Abdin, Sam Ade Jacobs, et al., Yifan Yang
arXiv Technical Report, 2024
Collaborative Preprints
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation (opens in new tab)
Yifan Zhou, Qihao Yang, Yan Li, et al., Yifan Yang, et al.
arXiv preprint arXiv:2607.08758, 2026 - WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces (opens in new tab)
Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, Yifan Yang, Dongsheng Li, Caihua Shan
arXiv preprint arXiv:2606.09426, 2026 - Covering Human Action Space for Computer Use: Data Synthesis and Benchmark (opens in new tab)
Miaosen Zhang, Xiaohan Zhao, Zhihong Tan, Zhou Huoshen, Yijia Fan, Yifan Yang, et al.
arXiv preprint arXiv:2605.12501, 2026 - EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents (opens in new tab)
Ruofei Ju, Xinrui Wang, Xin Ding, Yifan Yang, et al.
arXiv preprint arXiv:2605.10332, 2026 - MemCompiler: Compile, Don’t Inject—State-Conditioned Memory for Embodied Agents (opens in new tab)
Xin Ding, Xinrui Wang, Yifan Yang, Hao Wu, Shiqi Jiang, et al.
arXiv preprint arXiv:2605.07594, 2026
Collaborative Publications
- RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents (opens in new tab)
Jialiang Zhu, Gongrui Zhang, Xiaolong Ma, Lin Xu, Miaosen Zhang, Ruiqi Yang, Song Wang, Kai Qiu, Zhirong Wu, Qi Dai, Ruichun Ma, Bei Liu, Yifan Yang, et al.
ICML 2026 - AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation (opens in new tab)
Xin Ding, Jianyu Wei, Yifan Yang, et al.
ICML 2026 - AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information Minimization (opens in new tab)
Aoqi Wu, Junming Liu, Yuwei Zhang, Weiquan Huang, Liang Hu, Yifan Yang, et al.
Proceedings of the ACM Web Conference 2026 - A Comprehensive Ecosystem for Open-Domain Customized Video Generation (opens in new tab)
Jingxu Zhang, Yuqian Hong, Daneul Kim, Kai Qiu, Qi Dai, Jianmin Bao, Yifan Yang, et al.
ICASSP 2026 - Unified Medical Image Pre-training in Language-Guided Common Semantic Space (opens in new tab)
Xiaoxuan He, Yifan Yang, et al.
ECCV 2024 - StreamMind: Unlocking Full Frame-rate Streaming Video Dialogue through Event-gated Cognition (opens in new tab)
Xin Ding, Hao Wu, Yifan Yang, et al.
ICCV 2025 - Efficient and Adaptive Diffusion Model Inference through Lookup Tables on Mobile Devices (opens in new tab)
Qipeng Wang, Shiqi Jiang, Yifan Yang, et al.
IEEE Transactions on Mobile Computing, 2025 - Online Video Quality Enhancement with Spatial-Temporal Look-up Tables (opens in new tab)
Zefan Qu, Xinyang Jiang, Yifan Yang, et al.
ECCV 2024 - ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning (opens in new tab)
Rui Wang, Bohao Li, Yifan Yang, et al.
EMNLP 2025 - MageBench: Bridging Large Multimodal Models to Agents (opens in new tab)
Miaosen Zhang, Qi Dai, Yifan Yang, et al.
WACV 2025 - Reducio! Generating 1K Video within 16 Seconds Using Extremely Compressed Motion Latents (opens in new tab)
Rui Tian, Qi Dai, Yifan Yang, et al.
ICCV 2025 - Expand Heterogeneous Learning Systems with Selective Multi-Source Knowledge Fusion (opens in new tab)
Gaole Dai, Huatao Xu, Yifan Yang, Rui Tan, Mo Li
AAAI 2026 - Empowering Agentic Video Analytics Systems with Video Language Models (opens in new tab)
Yuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang, et al.
USENIX NSDI 2025 - DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation (opens in new tab)
Brian Nlong Zhao, Yifan Yang, et al.
ICLR 2025 - Understanding and Improving Training-free Loss-based Diffusion Guidance (opens in new tab)
Yifei Shen, Xinyang Jiang, Yifan Yang, et al.
NeurIPS 2024 - Online Video Super-resolution with Convolutional Kernel Bypass Grafts (opens in new tab)
Jun Xiao, Xinyang Jiang, Yifan Yang, et al.
IEEE Transactions on Multimedia, 2023
Contact
If you are interested in internship opportunities or research collaborations, feel free to reach out at
📧 yifanyang@microsoft.com