Research Intern, Visual Computing Group
Position: Research Intern
Job Type: Full-time Intern
Location: Beijing / Shanghai
Number of Openings: 1-2
Group Introduction
The Visual Computing Group at Microsoft Research Asia conducts cutting-edge research on multimodal foundation models, visual generation, and AI agents for creative and knowledge-work scenarios. Our mission is to advance the next generation of AI systems that can understand, reason, create, and collaborate with users through multimodal interactions.
Our current research spans:
- Multimodal reasoning and generation
- Agentic content creation and editing
- Visual generation and understanding
- Human-AI collaboration and AI-native productivity experiences
- Post-training and reinforcement learning for multimodal models
- Interactive design and controllable generation
- Data generation, evaluation, and model alignment
Our technologies are actively transferred to Microsoft Office, M365 Copilot, and future AI-native productivity experiences.
Researchers and interns in the group have opportunities to publish at top-tier conferences such as CVPR, ICCV, ECCV, NeurIPS, ICML, and ICLR, while also contributing to real-world product innovation.
Responsibilities
- Conduct cutting-edge research in multimodal AI, visual understanding and generation, and AI agents.
- Design, implement, and improve foundation models and agent systems to solve real-world problems.
- Build datasets, evaluation benchmarks, and experimentation pipelines for model development.
- Explore post-training, reinforcement learning, reasoning, and alignment techniques for multimodal models.
- Collaborate closely with research and product teams to translate research innovations into impactful user experiences.
- Publish research findings at leading academic venues and contribute to the broader research community.
Qualifications
Required
- Ability to devote most of your effort to the internship during the internship period.
- Experience in machine learning, deep learning, computer vision, natural language processing, or related fields.
- Familiarity with one or more of the following areas:
1.Large Language Models (LLMs)
2.Vision-Language Models (VLMs)
3.Diffusion and generative models
4.Reinforcement learning and post-training
5.AI agents and tool use
6.Multimodal reasoning and generation
- Strong problem-solving and research capabilities.
- Ability to read and communicate effectively in English.
- Strong communication and collaboration skills.
- Written approval from academic advisor.
Preferred
- Research experience demonstrated through publications, open-source contributions, competition achievements, or impactful projects.
- Experience with large-scale experimentation, evaluation, data curation, synthetic data generation, or model alignment.
- Experience developing agentic systems, multimodal applications, or productivity-related AI systems.
- Experience with distributed training, model deployment, or systems optimization.
Internship Duration
- Must obtain approval from academic advisor.
- Minimum internship duration: 3 months.
- Longer internship periods are strongly encouraged.