๐ŸผStalecollected in 56m

Chinese AI Models Overtake US in API Call Volume

Chinese AI Models Overtake US in API Call Volume
PostLinkedIn
๐ŸผRead original on Pandaily

๐Ÿ’กDiscover why Chinese models are dominating global API traffic and if they outperform your current stack.

โšก 30-Second TL;DR

What Changed

Chinese models lead OpenRouter API volume for 6 weeks

Why It Matters

This shift signals a growing preference for Chinese LLMs in global developer workflows, potentially challenging the dominance of US-based models.

What To Do Next

Integrate the DeepSeek-V4-Flash or MiniMax M3 APIs into your evaluation pipeline to compare their performance against GPT-4o.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขChinese models lead OpenRouter API volume for 6 weeks
  • โ€ขDeepSeek-V4-Flash is currently the most called model
  • โ€ขMiniMax M3 has achieved top-three global ranking

๐Ÿง  Deep Insight

Web-grounded analysis with 22 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขChinese AI models now command over 45% of OpenRouter's total traffic, a substantial increase from less than 2% in late 2024, indicating a significant shift in global developer adoption patterns.
  • โ€ขDeepSeek-V4-Flash is an efficiency-optimized Mixture-of-Experts (MoE) model, featuring 284 billion total parameters with only 13 billion activated per token, and supports a 1 million-token context window.
  • โ€ขMiniMax M3 distinguishes itself as the first open-weight model to integrate frontier-level coding capabilities, a 1-million-token context window, and native multimodal understanding (image and video input) within a single architecture.
  • โ€ขThe success and competitive pricing of Chinese models, exemplified by DeepSeek-V2 in May 2024, initiated a price war among major Chinese tech companies, leading to widespread reductions in AI model pricing.
  • โ€ขOpenRouter, a major AI model aggregation platform, processes over 36.1 trillion tokens weekly across hundreds of models, with Chinese models contributing 14.19 trillion tokens compared to 3.2 trillion from US models in the latest period.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature/MetricDeepSeek-V4-FlashMiniMax M3Claude Opus 4.8GPT-5.5Gemini 3.1 Pro
Input Price (per 1M tokens)$0.0983 - $0.14 (promo)$0.30 (promo) - $0.60$5.00$5.00$2.00
Output Price (per 1M tokens)$0.1966 - $0.28 (promo)$1.20 (promo) - $2.40$25.00$30.00$12.00
Context Window1M tokens1M tokensN/A (Opus 4.7 had 1M)1M-class tokens1M-class tokens
SWE-Bench ProN/A (V4 Pro Max 55.4%)59.0%69.2%Below M3Below M3
Terminal-Bench 2.1N/A (V4 Pro Max 67.9%)66.0%74.6% (Opus 4.8)N/AN/A
MultimodalityNative multimodal (text, images, video, audio)Native multimodal (text, image, video inputs)N/A (Opus 4.8 supports multimodal)N/A (GPT-5.4 Image 2 for image generation)N/A (Gemini 3.1 Pro supports audio/video multimodal depth)
ArchitectureMixture-of-Experts (MoE)MiniMax Sparse Attention (MSA)N/AN/AN/A
Open-Weight/SourceOpen-weight (MIT License)Open-weight (planned)Closed-sourceClosed-sourceClosed-source

๐Ÿ› ๏ธ Technical Deep Dive

  • DeepSeek-V4-Flash:

    • Architecture: Efficiency-optimized Mixture-of-Experts (MoE) model.
    • Parameters: 284 billion total parameters with 13 billion activated parameters.
    • Context Window: Supports a 1 million-token context window.
    • Attention Mechanism: Incorporates a Hybrid Attention Architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to significantly improve long-context efficiency, reducing computational overhead by approximately 50%.
    • Optimizations: Utilizes Manifold-Constrained Hyper-Connections (mHC) to strengthen residual connections and the Muon optimizer for faster convergence and training stability.
    • Efficiency: Achieves 10% of FLOPs and 7% of KV cache compared to DeepSeek-V3.2 for a 1M token context.
  • MiniMax M3:

    • Architecture: Built on proprietary MiniMax Sparse Attention (MSA).
    • Modality: Multimodal foundation model supporting text, image, and video inputs with text output.
    • Context Window: Features a 1 million-token context window.
    • Sparse Attention (MSA): Replaces full attention with KV-block selection, drastically cutting per-token compute at long context lengths (approximately 1/20 the cost of the previous generation at 1M tokens) and enabling substantially faster prefill and decode.
    • Performance: At 1M tokens, MSA delivers over 9x faster prefill and over 15x faster decoding compared to its predecessor, M2.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

The increasing market share of cost-effective Chinese AI models will intensify global price competition among AI providers.
Chinese models like DeepSeek-V4-Flash and MiniMax M3 offer significantly lower API pricing while demonstrating competitive performance, pressuring established US providers to re-evaluate their pricing strategies to remain competitive.
The trend towards open-weight and open-source models, particularly from Chinese developers, will accelerate innovation and adoption in agentic AI workflows.
MiniMax M3's open-weight nature, combined with its frontier capabilities and affordability, and DeepSeek's open-source approach, make these models highly attractive for developers building flexible and cost-efficient AI agents and applications.
The sustained dominance of Chinese AI models in API call volumes on platforms like OpenRouter will reshape perceptions of global AI leadership and influence investment flows.
The consistent lead of Chinese models in API usage on a major aggregation platform indicates a growing developer preference and technological maturity, potentially shifting investment and strategic focus towards Chinese AI innovation.

โณ Timeline

2021-12
MiniMax founded in Shanghai by former SenseTime executives.
2023-07
DeepSeek founded in Hangzhou, funded by High-Flyer.
2024-03
MiniMax raises $600 million in funding, reaching a $2.5 billion valuation.
2024-05
DeepSeek-V2 chatbot model released, gaining popularity in China and triggering a price war.
2025-01
DeepSeek-R1 model and eponymous chatbot launched, gaining international prominence and becoming a top downloaded app.
2026-01
MiniMax Group Inc. lists on the Hong Kong Stock Exchange.
2026-04-24
DeepSeek V4 Preview (including V4-Flash and V4-Pro) released with 1M context and MIT license.
2026-06-01
MiniMax M3 launched, featuring 1M context, native multimodality, and sparse attention architecture.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily โ†—