Xiaomi: The Token Price Butcher

💡Xiaomi's aggressive price cuts are shaking up the AI infrastructure market. Check if you can save on inference costs.
⚡ 30-Second TL;DR
What Changed
Aggressive reduction in token pricing for AI services
Why It Matters
Forces other AI providers to re-evaluate their pricing models to remain competitive in the mass-market developer space.
What To Do Next
Compare Xiaomi's new token pricing against your current LLM provider to see if switching can reduce your infrastructure overhead.
Key Points
- •Aggressive reduction in token pricing for AI services
- •Strategy focused on high-volume, low-cost AI accessibility
- •Direct competition with existing cloud and model providers on price
🧠 Deep Insight
Web-grounded analysis with 22 cited sources.
🔑 Enhanced Key Takeaways
- •Xiaomi has open-sourced its MiMo-V2.5 and MiMo-V2.5-Pro models under the permissive MIT License, making them suitable for commercial applications and positioning them as foundational infrastructure for the next generation of AI agents.
- •The MiMo-V2.5-Pro model, with 1.02 trillion parameters, demonstrates high token efficiency, requiring approximately 40-60% fewer tokens than comparable frontier models like Anthropic Claude Opus 4.6, Google Gemini 3.1 Pro, and OpenAI GPT-5.4 for agentic tasks.
- •Xiaomi's AI strategy is deeply integrated into its 'Human x Car x Home' smart ecosystem, aiming to extend AI beyond screens into physical world interactions through products like Xiaomi Miloco, which enables intelligent automation across smart home devices and electric vehicles.
- •The company has committed a substantial investment of over US$8.8 billion (60 billion yuan) in AI over the next three years, underscoring its ambition to develop frontier AI models and capabilities.
- •Xiaomi's pricing strategy includes not only aggressive API token cost reductions (up to 99% for cache hits) but also an overhauled 'Token Plan' subscription offering 5-8 times more usable tokens for the same price, and unified pricing across all context lengths.
📊 Competitor Analysis▸ Show
| Model / Provider | Input (Cache Hit) / Million Tokens | Input (Cache Miss) / Million Tokens | Output / Million Tokens | Key Features & Benchmarks |
|---|---|---|---|---|
| Xiaomi MiMo-V2.5-Pro | $0.0036 (RMB0.025) | $0.435 (RMB3) | $0.87 (RMB6) | 1.02T parameters, 1M context window, 40-60% more token efficient for agentic tasks than competitors. |
| Xiaomi MiMo-V2.5 | $0.0028 (RMB0.02) | $0.14 (RMB1) | $0.28 (RMB2) | 310B parameters, 1M context window, native omnimodal capabilities. |
| DeepSeek V4 Pro | $0.0036 (RMB0.025) | $0.435 (RMB3) | $0.87 (RMB6) | Xiaomi aligned pricing directly with DeepSeek V4 Pro. |
| DeepSeek V4 Flash | $0.0028 (RMB0.02) | $0.14 (RMB1) | $0.28 (RMB2) | Xiaomi aligned pricing directly with DeepSeek V4 Flash. |
| ByteDance Doubao (Seed-2.0-Pro) | N/A | $0.46 (RMB3.2) | $2.32 (RMB16) | Part of the 'price-cut camp' leveraging large ecosystems. |
| Kimi Moonshot V1 | N/A | $1.45 (RMB10) | $4.35 (RMB30) | Higher pricing, justified by long context and inference capabilities, achieved counter-trend rises in usage. |
| Anthropic Claude Opus 4.6 | N/A | N/A | $25.00 | Mentioned as requiring 40-60% more tokens for comparable agentic tasks than MiMo-V2.5-Pro. |
| Google Gemini 3.1 Pro | N/A | N/A | N/A | Mentioned as requiring 40-60% more tokens for comparable agentic tasks than MiMo-V2.5-Pro. |
| OpenAI GPT-5.4 | N/A | N/A | N/A | Mentioned as requiring 40-60% more tokens for comparable agentic tasks than MiMo-V2.5-Pro. |
🛠️ Technical Deep Dive
- MiMo-V2.5-Pro: Features 1.02 trillion total parameters with 42 billion active parameters, utilizing a Mixture-of-Experts (MoE) architecture. It supports a 1-million-token context window and employs a Hybrid Attention mechanism with a 7:1 ratio of Sliding Window Attention (SWA) to Global Attention (GA) for efficiency. A Multi-Token Prediction (MTP) module is integrated for faster generation, optimizing it for complex software engineering and long-horizon agentic tasks.
- MiMo-V2.5: A 310-billion-parameter model with native omnimodal capabilities, processing text, images, video, and audio within a single architecture. It also features a 1-million-token context window and delivers agentic performance comparable to its Pro sibling at approximately half the token cost.
- MiMo-V2-Flash: An open-weight Mixture-of-Experts (MoE) model with 309 billion total parameters and 15 billion active parameters. Its architecture includes a hybrid attention strategy (5:1 SWA to GA ratio) to reduce KV cache storage requirements by nearly 6x. It incorporates a learnable attention sink bias for coherence over sequences up to 256,000 tokens and uses Rollout Routing Replay (R3) to ensure training-inference consistency. A Multi-Token Prediction (MTP) module is embedded for self-speculative decoding, effectively tripling inference speeds.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (22)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Ifanr (爱范儿) ↗
