Kimi K3 Takes Over as K2.5 Retires

๐กKimi K2.5 is retiring soon, forcing developers to assess migration to a 2.8-trillion-parameter successor.
โก 30-Second TL;DR
What Changed
Kimi K2.5 will officially end service at the end of the month.
Why It Matters
The transition could affect application compatibility, latency, cost, and output quality for teams built on Kimi K2.5. It also marks a major model-generation shift for Moonshot AI and may intensify competition among large-scale LLM providers.
What To Do Next
Inventory every Kimi K2.5 dependency and run a regression test against Kimi K3 before the month-end retirement date.
Key Points
- โขKimi K2.5 will officially end service at the end of the month.
- โขKimi K2.5 is described as Moonshot AIโs first-generation trillion-parameter multimodal model.
- โขKimi K3 reportedly has 2.8 trillion parameters and has fully replaced K2.5.
- โขDevelopers using Kimi K2.5 may need to review migration plans before the retirement date.
๐ง Deep Insight
Background and context from public sources โ not the original article. 11 sources cited.
๐ Enhanced Key Takeaways
- โขMoonshot AI officially announced the retirement of Kimi K2.5 via their Weibo account, setting the final service termination date for August 31, 2026.
- โขKimi K2.5 had a notably short operational lifecycle of approximately eight months, having only launched in January 2026.
- โขKimi K3 is categorized as the world's first open-weights model in the 3-trillion-parameter class, specifically featuring 2.8 trillion parameters.
- โขThe transition to Kimi K3 includes a specific API pricing structure set at $0.30 per million tokens for cached input and $3.00 per million tokens for uncached input.
- โขMoonshot AI has faced intermittent consumer subscription pauses due to extreme infrastructure demand and computing capacity constraints following the K3 release.
๐ Competitor Analysisโธ Show
| Feature | Kimi K3 | DeepSeek V4 | Qwen 3.8 |
|---|---|---|---|
| Architecture | MoE (896 experts) | MoE | Dense/MoE Hybrid |
| Context Window | 1M Tokens | 512K Tokens | 1M Tokens |
| Primary Focus | Agentic/Long-horizon | Efficiency/Coding | General Purpose |
| Status | Open Weights | Open Weights | Open Weights |
๐ ๏ธ Technical Deep Dive
- Architecture: Mixture-of-Experts (MoE) framework utilizing 896 experts with 16 experts activated per token.
- Efficiency Innovation: Implements Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) to achieve 2.5x scaling efficiency over K2.5.
- Attention Mechanism: KDA replaces standard quadratic attention in specific layers to optimize long-context processing.
- Multimodality: Native support for integrated text, image, and video processing.
- Context Capacity: Native 1-million-token context window.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (11)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: cnBeta (Full RSS) โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.


