🔥Stalecollected in 2m

MiniMax M2.5 Tops Global LLM Calls 5 Weeks

MiniMax M2.5 Tops Global LLM Calls 5 Weeks
PostLinkedIn
🔥Read original on 36氪
#cost-efficiency#china-aiminimax-m2.5minimaxm2.5

💡Chinese LLM leads global usage w/ 10x cost edge via tech + energy – must-eval for prod scale

⚡ 30-Second TL;DR

What Changed

MiniMax M2.5 #1 in global model calls for 5 consecutive weeks

Why It Matters

Demonstrates China's rising AI dominance via cost leadership, pressuring Western providers to cut prices. Enables broader global adoption of high-performance LLMs in production. Shifts market toward efficiency-focused models.

What To Do Next

Benchmark MiniMax M2.5 API for your inference workloads to achieve 10x cost savings.

Who should care:Developers & AI Engineers

Key Points

  • MiniMax M2.5 #1 in global model calls for 5 consecutive weeks
  • Up to 10x cheaper than overseas models for equivalent performance
  • Architecture innovation cuts inference costs via fewer tokens
  • China's low electricity prices (70-80% of compute costs) boost edge

🧠 Deep Insight

Background and context from public sources — not the original article. 10 sources cited.

🔑 Enhanced Key Takeaways

  • MiniMax M2.5 utilizes a Mixture-of-Experts (MoE) architecture with 230 billion total parameters, activating only ~10 billion per request, which enables 20x computational efficiency and high-speed inference.
  • The model incorporates 'Lightning Attention' to replace quadratic complexity with linear scaling, allowing for efficient processing of its 200,000-token context window.
  • M2.5 was trained using reinforcement learning across hundreds of thousands of real-world environments to optimize task decomposition and reasoning, achieving SOTA performance on benchmarks like SWE-Bench Verified (80.2%).
📊 Competitor Analysis▸ Show
FeatureMiniMax M2.5Claude Opus 4.6GPT-5 Series
Architecture230B MoE (10B active)Proprietary Dense/MoEProprietary Dense/MoE
Output Pricing~$1.20 - $2.40/1M tokens~$25/1M tokens~$60/1M tokens
Throughput50-100 tokens/sec~33 tokens/sec~40-50 tokens/sec
SWE-Bench Verified80.2%80.8%80.0%

🛠️ Technical Deep Dive

  • Architecture: Mixture-of-Experts (MoE) with 230B total parameters, ~10B active parameters per inference.
  • Attention Mechanism: Lightning Attention, which reformulates attention as a streaming process to achieve linear scaling instead of quadratic.
  • Training: Extensive reinforcement learning (RL) focused on per-token process rewards to improve task decomposition and reasoning efficiency.
  • Context Window: 200,000 tokens.
  • Inference Variants: Standard (50 tokens/sec) and Lightning (100 tokens/sec).
  • Coding Capability: Native 'spec-writing' behavior where the model plans architecture and structure before generating code.

🔮 Future ImplicationsAI analysis grounded in cited sources

Agentic workflows will shift from being cost-prohibitive to standard enterprise practice.
The drastic reduction in inference costs allows for 'always-on' autonomous agents that were previously economically unviable.
Open-weights frontier models will erode the market share of proprietary closed-source models.
Performance parity with top-tier proprietary models combined with significantly lower costs and open availability lowers the barrier for enterprise adoption.

Timeline

2021-12
MiniMax founded by former SenseTime researchers.
2024-03
Alibaba leads $600 million financing round; company valuation reaches $2.5 billion.
2026-01
MiniMax completes initial public offering on the Hong Kong Stock Exchange.
2026-02
MiniMax releases M2.5 model series.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.