Kimi K3 Seeks 30% Share on U.S. Clouds

💡Kimi K3 could become the first Chinese model offered through U.S. hyperscaler clouds.
⚡ 30-Second TL;DR
What Changed
Moonshot AI is discussing Kimi K3 cloud deployments with Microsoft, Amazon, and Google.
Why It Matters
A successful deal could improve Kimi K3’s access to global enterprise customers while giving U.S. cloud providers a commercial route to offer a Chinese model. Developers should also consider pricing, data-governance, and geopolitical constraints before adopting it in production.
What To Do Next
Model the total cost of serving Kimi K3 under a 30% hyperscaler revenue share and compare it with your current inference provider.
Key Points
- •Moonshot AI is discussing Kimi K3 cloud deployments with Microsoft, Amazon, and Google.
- •The proposed revenue share for hyperscalers could reach 30%.
- •Kimi K3 would reportedly be the first Chinese model deployed on U.S. hyperscaler clouds under this type of arrangement.
🧠 Deep Insight
Background and context from public sources — not the original article. 10 sources cited.
🔑 Enhanced Key Takeaways
- •Kimi K3 utilizes a Mixture-of-Experts (MoE) architecture featuring 896 experts, with 104 billion parameters active during inference.
- •The model incorporates proprietary technical innovations identified as 'Kimi Delta Attention' and 'Stable LatentMoE' to manage its 2.8 trillion parameter scale.
- •Moonshot AI has already validated its revenue-sharing model through existing partnerships with smaller entities like Chinasoft International and inference providers such as Together AI.
- •The negotiations face significant friction regarding technical auditing of token usage and data access provisions, which remain unresolved as of September 2026.
- •Moonshot AI is currently navigating U.S. political scrutiny concerning its historical chip procurement and development practices, which complicates the approval process for hyperscaler integration.
📊 Competitor Analysis▸ Show
| Feature | Kimi K3 | GPT-4o | Claude 3.5 Sonnet |
|---|---|---|---|
| Parameters | 2.8T (MoE) | Undisclosed | Undisclosed |
| Context Window | 1M Tokens | 128K Tokens | 200K Tokens |
| Architecture | MoE (896 experts) | MoE | Dense/Hybrid |
| Deployment | Open-weight/Cloud | Closed API | Closed API |
🛠️ Technical Deep Dive
- Model Size: 2.8 trillion parameters total.
- Architecture: Mixture-of-Experts (MoE) with 896 experts.
- Active Parameters: 104 billion parameters active per inference pass.
- Context Window: 1 million tokens.
- Proprietary Modules: Kimi Delta Attention and Stable LatentMoE.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (10)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.


