Chinese AI Models Overtake US in API Call Volume

๐กDiscover why Chinese models are dominating global API traffic and if they outperform your current stack.
โก 30-Second TL;DR
What Changed
Chinese models lead OpenRouter API volume for 6 weeks
Why It Matters
This shift signals a growing preference for Chinese LLMs in global developer workflows, potentially challenging the dominance of US-based models.
What To Do Next
Integrate the DeepSeek-V4-Flash or MiniMax M3 APIs into your evaluation pipeline to compare their performance against GPT-4o.
Key Points
- โขChinese models lead OpenRouter API volume for 6 weeks
- โขDeepSeek-V4-Flash is currently the most called model
- โขMiniMax M3 has achieved top-three global ranking
๐ง Deep Insight
Web-grounded analysis with 22 cited sources.
๐ Enhanced Key Takeaways
- โขChinese AI models now command over 45% of OpenRouter's total traffic, a substantial increase from less than 2% in late 2024, indicating a significant shift in global developer adoption patterns.
- โขDeepSeek-V4-Flash is an efficiency-optimized Mixture-of-Experts (MoE) model, featuring 284 billion total parameters with only 13 billion activated per token, and supports a 1 million-token context window.
- โขMiniMax M3 distinguishes itself as the first open-weight model to integrate frontier-level coding capabilities, a 1-million-token context window, and native multimodal understanding (image and video input) within a single architecture.
- โขThe success and competitive pricing of Chinese models, exemplified by DeepSeek-V2 in May 2024, initiated a price war among major Chinese tech companies, leading to widespread reductions in AI model pricing.
- โขOpenRouter, a major AI model aggregation platform, processes over 36.1 trillion tokens weekly across hundreds of models, with Chinese models contributing 14.19 trillion tokens compared to 3.2 trillion from US models in the latest period.
๐ Competitor Analysisโธ Show
| Feature/Metric | DeepSeek-V4-Flash | MiniMax M3 | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| Input Price (per 1M tokens) | $0.0983 - $0.14 (promo) | $0.30 (promo) - $0.60 | $5.00 | $5.00 | $2.00 |
| Output Price (per 1M tokens) | $0.1966 - $0.28 (promo) | $1.20 (promo) - $2.40 | $25.00 | $30.00 | $12.00 |
| Context Window | 1M tokens | 1M tokens | N/A (Opus 4.7 had 1M) | 1M-class tokens | 1M-class tokens |
| SWE-Bench Pro | N/A (V4 Pro Max 55.4%) | 59.0% | 69.2% | Below M3 | Below M3 |
| Terminal-Bench 2.1 | N/A (V4 Pro Max 67.9%) | 66.0% | 74.6% (Opus 4.8) | N/A | N/A |
| Multimodality | Native multimodal (text, images, video, audio) | Native multimodal (text, image, video inputs) | N/A (Opus 4.8 supports multimodal) | N/A (GPT-5.4 Image 2 for image generation) | N/A (Gemini 3.1 Pro supports audio/video multimodal depth) |
| Architecture | Mixture-of-Experts (MoE) | MiniMax Sparse Attention (MSA) | N/A | N/A | N/A |
| Open-Weight/Source | Open-weight (MIT License) | Open-weight (planned) | Closed-source | Closed-source | Closed-source |
๐ ๏ธ Technical Deep Dive
-
DeepSeek-V4-Flash:
- Architecture: Efficiency-optimized Mixture-of-Experts (MoE) model.
- Parameters: 284 billion total parameters with 13 billion activated parameters.
- Context Window: Supports a 1 million-token context window.
- Attention Mechanism: Incorporates a Hybrid Attention Architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to significantly improve long-context efficiency, reducing computational overhead by approximately 50%.
- Optimizations: Utilizes Manifold-Constrained Hyper-Connections (mHC) to strengthen residual connections and the Muon optimizer for faster convergence and training stability.
- Efficiency: Achieves 10% of FLOPs and 7% of KV cache compared to DeepSeek-V3.2 for a 1M token context.
-
MiniMax M3:
- Architecture: Built on proprietary MiniMax Sparse Attention (MSA).
- Modality: Multimodal foundation model supporting text, image, and video inputs with text output.
- Context Window: Features a 1 million-token context window.
- Sparse Attention (MSA): Replaces full attention with KV-block selection, drastically cutting per-token compute at long context lengths (approximately 1/20 the cost of the previous generation at 1M tokens) and enabling substantially faster prefill and decode.
- Performance: At 1M tokens, MSA delivers over 9x faster prefill and over 15x faster decoding compared to its predecessor, M2.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (22)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily โ
