One Model, Many Modalities: Building Unified Foundation Models Across Text, Speech, Video and Music
Session Overview
We are witnessing a fundamental shift in the artificial intelligence landscape: the convergence of single-modality architectures into unified, natively multimodal foundation models. As generative AI advances beyond isolated text processing, the future belongs to systems that seamlessly understand, reason, and generate across text, speech, video, and music within a single cohesive framework.
This keynote explores the driving forces behind this industry transformation, breaks down the latest architectural breakthroughs enabling true cross-modal unity, and examines the evolving open-source model ecosystem. Drawing directly from MiniMax's frontier research and practical deployments, Cherie Shi reveals MiniMax's approach to building high-performance multimodal foundation models that push the boundaries of real-time human-AI interaction and global creativity.
What This Session Covers
- Industry change driving the shift toward unified multimodal AI
- The latest technological advances in cross-modal architecture
- The current open source model landscape
- MiniMax's latest model releases and multimodal approach
Three Key Takeaways
- The Unified Modality Shift — How the industry and technical paradigms are rapidly transitioning from fragmented, specialized models to natively unified architectures across text, audio, video, and music.
- Open Source vs. Frontier Innovations — An inside look at the current open-source multimodal landscape, highlighting key technological gaps, emerging standards, and where breakthrough value lies.
- The MiniMax Multimodal Playbook — Concrete insights into MiniMax's latest model releases and engineering approach to scaling, training, and deploying cross-modal foundation models for high-impact real-world applications.
Speaker
Cherie Shi is a Regional Manager at MiniMax. She partners with key clients and strategic partners worldwide to unlock the full potential of large language models and multimodal models, and collaborates closely with research and product teams on model evaluation and harness optimization.
As a founding team member of MiniMax, Cherie previously served as VP of Marketing and Strategy, leading go-to-market strategies, multimodal model launches, corporate strategy, and business analytics. She holds a bachelor's degree from Tsinghua University and brings eight years of experience from leading AI research labs.

