⚛️Stalecollected in 3h

MiniMax M3 Hands-on: Multimodal Reasoning Test

MiniMax M3 Hands-on: Multimodal Reasoning Test
PostLinkedIn
⚛️Read original on 量子位

💡Can a Chinese multimodal model handle complex Nvidia slides? See how MiniMax M3 performs in real-world visual tests.

⚡ 30-Second TL;DR

What Changed

Tested MiniMax M3 on complex Nvidia presentation slides

Why It Matters

Shows the rapid progress of Chinese multimodal models in competing with global SOTA benchmarks. This indicates a narrowing gap in visual understanding tasks.

What To Do Next

Evaluate MiniMax M3's API for your specific visual extraction workflows to compare against GPT-4o.

Who should care:Developers & AI Engineers

Key Points

  • Tested MiniMax M3 on complex Nvidia presentation slides
  • Evaluated visual recognition and reasoning performance
  • Demonstrated capability in handling dense visual data

🧠 Deep Insight

Web-grounded analysis with 12 cited sources.

🔑 Enhanced Key Takeaways

  • MiniMax M3, officially released on June 1, 2026, is positioned as the first open-weight model to integrate frontier coding performance, a 1-million-token context window, and native multimodal capabilities, including image and video input, and the ability to operate a desktop computer.
  • The model incorporates a novel 'MiniMax Sparse Attention (MSA)' architecture, which drastically improves computational efficiency, reducing per-token compute at 1M context to 1/20th of the previous generation and enabling 9x faster prefill and 15x faster decoding.
  • M3 demonstrates strong performance on key benchmarks, surpassing GPT-5.5 and Gemini 3.1 Pro on SWE-Bench Pro (59.0%) and outperforming Claude Opus 4.7 on SVG-Bench and BrowseComp (83.5%), while being offered at a significantly lower cost.
  • The model was developed with native multimodality from 'step zero,' meaning text, images, and video were trained together from the outset using interleaved data, rather than adding modalities post-training.
  • MiniMax M3 is specifically designed for demanding applications such as long-horizon agent tasks, comprehensive codebase analysis, and in-depth long-video understanding, leveraging its extensive context window and multimodal processing to facilitate complex, multi-step workflows.
📊 Competitor Analysis▸ Show
Feature/MetricMiniMax M3 (Open-Weight)Claude Opus 4.8 (Closed-Source)GPT-5.5 (Closed-Source)Gemini 3.1 Pro (Closed-Source)
Release DateJune 1, 2026May 25, 2026N/A (Pre-M3)N/A (Pre-M3)
Context Window1M tokensN/A (typically large)1M-class1M-class
MultimodalityNative (Text, Image, Video input, Desktop operation)Yes (Text, Image, Video)Yes (Omnimodal)Yes (Audio/Video depth)
SWE-Bench Pro59.0%69.2%58.6%54.2%
Terminal-Bench 2.166.0%74.6% (Opus 4.8)N/AN/A
BrowseComp83.5%79.3% (Opus 4.7)N/AN/A
SVG-BenchSurpasses Opus 4.7N/AN/AN/A
Pricing (Input/Output per 1M tokens, promo)$0.30 / $1.20Significantly higherSignificantly higherSignificantly higher
Key ArchitectureMSA (Sparse Attention)N/AN/AN/A

🛠️ Technical Deep Dive

  • Architecture: MiniMax M3 is built on a proprietary MiniMax Sparse Attention (MSA) architecture, which is a clean and scalable sparse attention mechanism. It is also described as a sparse Mixture-of-Experts (MoE) foundation model.
  • Context Window: Supports an ultra-long context window of up to 1 million tokens, with a guaranteed minimum of 512K tokens.
  • Multimodality: Natively multimodal, supporting image and video input, and capable of operating a desktop computer. This capability was trained from 'step zero' using interleaved data, where text, images, and video are processed together from the beginning.
  • Computational Efficiency: The MSA architecture reduces per-token compute at a 1 million token context length to just 1/20th of the previous generation model.
  • Speed Improvements: Achieves more than 9x faster prefilling and more than 15x faster decoding at a 1 million token context length compared to previous models.
  • Training Data: The entire data pipeline was rebuilt for interleaved formats, and training data was scaled to the order of 100 trillion tokens.

🔮 Future ImplicationsAI analysis grounded in cited sources

MiniMax M3 will accelerate the adoption of open-weight models for complex agentic and coding workflows.
Its combination of frontier performance, large context, native multimodality, and competitive pricing makes it a compelling alternative to closed-source models for developers.
The MiniMax Sparse Attention (MSA) architecture will influence future large language model designs, particularly for long-context processing.
Its demonstrated efficiency gains in prefill and decoding at 1M tokens address a major challenge in scaling context windows without quadratic computational complexity.
MiniMax will strengthen its position as a global leader in multimodal AI, particularly in the agentic AI space.
M3's capabilities in desktop operation, long-horizon tasks, and strong benchmarks in agentic evaluations position it well for developing advanced AI agents.

Timeline

2021-12
MiniMax founded by computer vision researchers from SenseTime.
2022-10
MiniMax launched Glow, an AI character app.
2024-03
Alibaba Group led a $600 million financing round for MiniMax, valuing it at $2.5 billion.
2025-01
MiniMax unveiled the MiniMax-01 LLM product line, including MiniMax-Text-01 and MiniMax-VL-01 with visual capabilities.
2026-01
MiniMax Group Inc. listed on the Hong Kong Stock Exchange.
2026-06
MiniMax M3 officially released, featuring frontier coding, 1M context, and native multimodality.

📎 Sources (12)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. minimax.io
  2. techtimes.com
  3. marktechpost.com
  4. kingy.ai
  5. siliconflow.com
  6. lushbinary.com
  7. qubrid.com
  8. gigazine.net
  9. venturebeat.com
  10. lushbinary.com
  11. ollama.com
  12. aimlapi.com
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位