๐Ÿ’ผStalecollected in 2m

MiniMax teases M3 model with 15.6X faster long-context speed

MiniMax teases M3 model with 15.6X faster long-context speed
PostLinkedIn
๐Ÿ’ผRead original on VentureBeat

๐Ÿ’กDiscover how MiniMax's new sparse attention mechanism achieves a 15.6X speed boost for long-context AI agents.

โšก 30-Second TL;DR

What Changed

MiniMax M3 will feature a custom sub-quadratic sparse attention framework.

Why It Matters

This breakthrough in decoding speed significantly lowers the barrier for deploying complex, long-context AI agents in enterprise production environments.

What To Do Next

Review the MiniMax M2 technical report to understand their MoE routing strategies for optimizing your own model fine-tuning workflows.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขMiniMax M3 will feature a custom sub-quadratic sparse attention framework.
  • โ€ขThe M2 series uses a Mixture-of-Experts (MoE) architecture with 229.9B total parameters.
  • โ€ขM2 achieves efficiency by activating only 9.8B parameters per token across 256 experts.
  • โ€ขThe new M3 architecture aims to make ultra-long-context AI agent deployment economically viable.

๐Ÿง  Deep Insight

Web-grounded analysis with 13 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขMiniMax, a Chinese AI startup founded in late 2021 or early 2022 by former SenseTime executives, went public on the Hong Kong Stock Exchange in January 2026.
  • โ€ขPrior to its IPO, MiniMax successfully raised $1.5 billion across seven funding rounds, with Alibaba emerging as the largest external shareholder and Hillhouse Capital as the earliest investor.
  • โ€ขThe M2 series, including the M2.7 model, has demonstrated strong performance in agentic and coding benchmarks, with M2.7 scoring 86.2% on PinchBench and ranking highly among open-source models globally.
  • โ€ขBeyond the 15.6x decoding speedup, MiniMax's M3 model is also projected to achieve a 9.7x faster prefill speed specifically at 1-million-token context lengths compared to M2.
  • โ€ขThere are indications from MiniMax's head of engineering that the M3 model, along with other tools like Mavis and Teams, may be released as open-source.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature/ModelMiniMax M2.7MiniMax M1-80kClaude Opus 4.6GPT-5.4GLM-5Qwen3.5-plusGPT-oss 120B
ArchitectureMoE (229.9B total, 9.8B active)MoE, Hybrid-Attention (456B total, 45.9B active)---MoEMoE (117B total, 5.1B active)
Long Context-1M tokens native support-----
PinchBench (Agent Tasks)86.2% (5th overall)->86.2% (implied higher than M2.7)86.4%86.4%85.8%-
Kilo Bench (Autonomous Coding)47% (2nd overall)---->47% (implied higher than M2.7)-
MMLU85.0% (M2.5)-88.5%-86.7%85.5%91.8%
Cost (per M tokens)$0.30 input / $1.20 output-Much higher than M2.7 (implied)Much higher than M2.7 (implied)---

๐Ÿ› ๏ธ Technical Deep Dive

  • The M2 series employs a sparse Mixture-of-Experts (MoE) decoder-only Transformer architecture, featuring 229.9 billion total parameters with only 9.8 billion parameters activated per token across 256 experts.
  • MiniMax optimized routing in M2 by implementing sigmoid gating paired with learnable, expert-specific bias terms.
  • While M2 development rigorously tested sub-quadratic shortcuts, they were initially discarded due to concerns about crippling the model's 'multi-hop reasoning' capabilities, leading to the adoption of full quadratic attention to maintain intelligence.
  • The M3 model's custom sub-quadratic sparse attention framework is designed to perform attention directly on the real KV cache, rather than operating in compressed dimensions, which is a distinction from approaches like DeepSeek V4's CSA.
  • MiniMax's M2.1, an open-source model, was developed with post-training optimization and utilizes an internal framework called 'Forge' specifically for agent-centric scenarios.
  • The training for M2.1 involved agent-driven automated data pipelines that leveraged raw GitHub data to generate diverse and verifiable software engineering datasets and environments.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

MiniMax's M3 will significantly lower the barrier to entry for ultra-long-context AI agent deployment.
The claimed 15.6x decode speedup at 1M-token contexts would make such inference economically viable, addressing current latency and cost issues that deter enterprise customers.
The AI industry will see accelerated adoption and development of sparse attention mechanisms in large language models.
MiniMax's public preview of M3's sparse attention signals its production readiness, which will likely pressure competitors like Anthropic, Google DeepMind, and OpenAI to advance their own efficient-attention roadmaps.
MiniMax will strengthen its position in the global open-source LLM ecosystem.
Hints from MiniMax's head of engineering suggest M3, along with other tools, might be open-sourced, which would increase its influence and adoption among developers and researchers.

โณ Timeline

2021-12
MiniMax founded by former SenseTime researchers.
2022-10
MiniMax launches its first product, Glow.
2023-03
MiniMax closes its third funding round, raising $260 million at a $1.157 billion valuation.
2024-03
Alibaba Group leads a $600 million financing round, valuing MiniMax at $2.5 billion.
2025-07
MiniMax secures nearly $300 million in a funding round, reaching a post-money valuation of approximately $4 billion, with state-owned entity investment.
2026-01-09
MiniMax holds its Initial Public Offering (IPO) on the Hong Kong Stock Exchange.
2026-03-18
MiniMax-M2.7 is listed as the most recent M-series text model in platform release notes.

๐Ÿ“Ž Sources (13)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. htx.com
  2. wikipedia.org
  3. techflowpost.com
  4. github.com
  5. reddit.com
  6. aiweekly.co
  7. reddit.com
  8. venturebeat.com
  9. siliconflow.com
  10. huggingface.co
  11. onyx.app
  12. youtube.com
  13. minimax.io
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: VentureBeat โ†—