Chinese AI Models Overtake US in API Call Volume

💡Discover why Chinese models are dominating global API traffic and if they outperform your current stack.
⚡ 30-Second TL;DR
What Changed
Chinese models lead OpenRouter API volume for 6 weeks
Why It Matters
This shift signals a growing preference for Chinese LLMs in global developer workflows, potentially challenging the dominance of US-based models.
What To Do Next
Integrate the DeepSeek-V4-Flash or MiniMax M3 APIs into your evaluation pipeline to compare their performance against GPT-4o.
Key Points
- •Chinese models lead OpenRouter API volume for 6 weeks
- •DeepSeek-V4-Flash is currently the most called model
- •MiniMax M3 has achieved top-three global ranking
🧠 Deep Insight
Background and context from public sources — not the original article. 22 sources cited.
🔑 Enhanced Key Takeaways
- •Chinese AI models now command over 45% of OpenRouter's total traffic, a substantial increase from less than 2% in late 2024, indicating a significant shift in global developer adoption patterns.
- •DeepSeek-V4-Flash is an efficiency-optimized Mixture-of-Experts (MoE) model, featuring 284 billion total parameters with only 13 billion activated per token, and supports a 1 million-token context window.
- •MiniMax M3 distinguishes itself as the first open-weight model to integrate frontier-level coding capabilities, a 1-million-token context window, and native multimodal understanding (image and video input) within a single architecture.
- •The success and competitive pricing of Chinese models, exemplified by DeepSeek-V2 in May 2024, initiated a price war among major Chinese tech companies, leading to widespread reductions in AI model pricing.
- •OpenRouter, a major AI model aggregation platform, processes over 36.1 trillion tokens weekly across hundreds of models, with Chinese models contributing 14.19 trillion tokens compared to 3.2 trillion from US models in the latest period.
📊 Competitor Analysis▸ Show
| Feature/Metric | DeepSeek-V4-Flash | MiniMax M3 | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| Input Price (per 1M tokens) | $0.0983 - $0.14 (promo) | $0.30 (promo) - $0.60 | $5.00 | $5.00 | $2.00 |
| Output Price (per 1M tokens) | $0.1966 - $0.28 (promo) | $1.20 (promo) - $2.40 | $25.00 | $30.00 | $12.00 |
| Context Window | 1M tokens | 1M tokens | N/A (Opus 4.7 had 1M) | 1M-class tokens | 1M-class tokens |
| SWE-Bench Pro | N/A (V4 Pro Max 55.4%) | 59.0% | 69.2% | Below M3 | Below M3 |
| Terminal-Bench 2.1 | N/A (V4 Pro Max 67.9%) | 66.0% | 74.6% (Opus 4.8) | N/A | N/A |
| Multimodality | Native multimodal (text, images, video, audio) | Native multimodal (text, image, video inputs) | N/A (Opus 4.8 supports multimodal) | N/A (GPT-5.4 Image 2 for image generation) | N/A (Gemini 3.1 Pro supports audio/video multimodal depth) |
| Architecture | Mixture-of-Experts (MoE) | MiniMax Sparse Attention (MSA) | N/A | N/A | N/A |
| Open-Weight/Source | Open-weight (MIT License) | Open-weight (planned) | Closed-source | Closed-source | Closed-source |
🛠️ Technical Deep Dive
-
DeepSeek-V4-Flash:
- Architecture: Efficiency-optimized Mixture-of-Experts (MoE) model.
- Parameters: 284 billion total parameters with 13 billion activated parameters.
- Context Window: Supports a 1 million-token context window.
- Attention Mechanism: Incorporates a Hybrid Attention Architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to significantly improve long-context efficiency, reducing computational overhead by approximately 50%.
- Optimizations: Utilizes Manifold-Constrained Hyper-Connections (mHC) to strengthen residual connections and the Muon optimizer for faster convergence and training stability.
- Efficiency: Achieves 10% of FLOPs and 7% of KV cache compared to DeepSeek-V3.2 for a 1M token context.
-
MiniMax M3:
- Architecture: Built on proprietary MiniMax Sparse Attention (MSA).
- Modality: Multimodal foundation model supporting text, image, and video inputs with text output.
- Context Window: Features a 1 million-token context window.
- Sparse Attention (MSA): Replaces full attention with KV-block selection, drastically cutting per-token compute at long context lengths (approximately 1/20 the cost of the previous generation at 1M tokens) and enabling substantially faster prefill and decode.
- Performance: At 1M tokens, MSA delivers over 9x faster prefill and over 15x faster decoding compared to its predecessor, M2.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (22)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.

