MiniMax M3 Hands-on: Multimodal Reasoning Test

💡Can a Chinese multimodal model handle complex Nvidia slides? See how MiniMax M3 performs in real-world visual tests.
⚡ 30-Second TL;DR
What Changed
Tested MiniMax M3 on complex Nvidia presentation slides
Why It Matters
Shows the rapid progress of Chinese multimodal models in competing with global SOTA benchmarks. This indicates a narrowing gap in visual understanding tasks.
What To Do Next
Evaluate MiniMax M3's API for your specific visual extraction workflows to compare against GPT-4o.
Key Points
- •Tested MiniMax M3 on complex Nvidia presentation slides
- •Evaluated visual recognition and reasoning performance
- •Demonstrated capability in handling dense visual data
🧠 Deep Insight
Web-grounded analysis with 12 cited sources.
🔑 Enhanced Key Takeaways
- •MiniMax M3, officially released on June 1, 2026, is positioned as the first open-weight model to integrate frontier coding performance, a 1-million-token context window, and native multimodal capabilities, including image and video input, and the ability to operate a desktop computer.
- •The model incorporates a novel 'MiniMax Sparse Attention (MSA)' architecture, which drastically improves computational efficiency, reducing per-token compute at 1M context to 1/20th of the previous generation and enabling 9x faster prefill and 15x faster decoding.
- •M3 demonstrates strong performance on key benchmarks, surpassing GPT-5.5 and Gemini 3.1 Pro on SWE-Bench Pro (59.0%) and outperforming Claude Opus 4.7 on SVG-Bench and BrowseComp (83.5%), while being offered at a significantly lower cost.
- •The model was developed with native multimodality from 'step zero,' meaning text, images, and video were trained together from the outset using interleaved data, rather than adding modalities post-training.
- •MiniMax M3 is specifically designed for demanding applications such as long-horizon agent tasks, comprehensive codebase analysis, and in-depth long-video understanding, leveraging its extensive context window and multimodal processing to facilitate complex, multi-step workflows.
📊 Competitor Analysis▸ Show
| Feature/Metric | MiniMax M3 (Open-Weight) | Claude Opus 4.8 (Closed-Source) | GPT-5.5 (Closed-Source) | Gemini 3.1 Pro (Closed-Source) |
|---|---|---|---|---|
| Release Date | June 1, 2026 | May 25, 2026 | N/A (Pre-M3) | N/A (Pre-M3) |
| Context Window | 1M tokens | N/A (typically large) | 1M-class | 1M-class |
| Multimodality | Native (Text, Image, Video input, Desktop operation) | Yes (Text, Image, Video) | Yes (Omnimodal) | Yes (Audio/Video depth) |
| SWE-Bench Pro | 59.0% | 69.2% | 58.6% | 54.2% |
| Terminal-Bench 2.1 | 66.0% | 74.6% (Opus 4.8) | N/A | N/A |
| BrowseComp | 83.5% | 79.3% (Opus 4.7) | N/A | N/A |
| SVG-Bench | Surpasses Opus 4.7 | N/A | N/A | N/A |
| Pricing (Input/Output per 1M tokens, promo) | $0.30 / $1.20 | Significantly higher | Significantly higher | Significantly higher |
| Key Architecture | MSA (Sparse Attention) | N/A | N/A | N/A |
🛠️ Technical Deep Dive
- Architecture: MiniMax M3 is built on a proprietary MiniMax Sparse Attention (MSA) architecture, which is a clean and scalable sparse attention mechanism. It is also described as a sparse Mixture-of-Experts (MoE) foundation model.
- Context Window: Supports an ultra-long context window of up to 1 million tokens, with a guaranteed minimum of 512K tokens.
- Multimodality: Natively multimodal, supporting image and video input, and capable of operating a desktop computer. This capability was trained from 'step zero' using interleaved data, where text, images, and video are processed together from the beginning.
- Computational Efficiency: The MSA architecture reduces per-token compute at a 1 million token context length to just 1/20th of the previous generation model.
- Speed Improvements: Achieves more than 9x faster prefilling and more than 15x faster decoding at a 1 million token context length compared to previous models.
- Training Data: The entire data pipeline was rebuilt for interleaved formats, and training data was scaled to the order of 100 trillion tokens.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (12)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗
