Kimi K2.8 Preview Brings Million-Token Context

Kimi adds 1M context and adjustable reasoning without changing its coding API ID.
30-Second TL;DR
What Changed
K2.8 Preview is live for Kimi Code and Work users.
Why It Matters
A large context window and adjustable reasoning could benefit repository-scale coding and long-running agent tasks. The unchanged API identifier may reduce migration work for existing developers.
What To Do Next
Run your largest repository workflows against kimi-for-coding and measure context retention, latency, and tool-call reliability.
Key Points
- •K2.8 Preview is live for Kimi Code and Work users.
- •All tiers reportedly receive a 1M-token context window.
- •Users can adjust reasoning, while the coding API ID remains unchanged.
Deep Insight
Background and context from public sources — not the original article. 10 sources cited.
Enhanced Key Takeaways
- •Kimi K2.8 Preview serves as an upgrade from Kimi K2.7 Code, expanding the context window fourfold from 256K tokens to 1,048,576 tokens.
- •Adjustable reasoning controls offer granular settings across 'low', 'high', and 'max', with 'max' configured as the default reasoning mode.
- •Under a new architectural routing rule, disabling thinking on K3 or K2.8 queries automatically diverts execution to K2.8's non-thinking pipeline.
- •The model natively accepts multimodal inputs, supporting image and video processing in addition to large-scale multi-file codebase analysis.
- •Developer endpoints achieve inference throughput speeds reaching up to approximately 260 tokens per second for high-efficiency agentic workflows.
Competitor Analysis
- Context Window
- 1,048,576 tokens (All tiers)
- Reasoning Modes
- Configurable (
low,high,max) - Multimodal Input
- Text, Code, Image, Video
- Primary Role / Positioning
- High-throughput agentic daily driver approaching K3 quality
- Context Window
- 256,000 tokens
- Reasoning Modes
- Fixed / Limited
- Multimodal Input
- Text, Code
- Primary Role / Positioning
- Previous-generation coding model with higher reasoning overhead
- Context Window
- 1,048,576 tokens (Enterprise)
- Reasoning Modes
- Extended thinking
- Multimodal Input
- Text, Code, Multimodal
- Primary Role / Positioning
- 2.8T MoE flagship frontier model
- Context Window
- 200,000 tokens
- Reasoning Modes
- Extended thinking
- Multimodal Input
- Text, Code, Image
- Primary Role / Positioning
- Western benchmark for terminal-based agentic software engineering
| Model / Tool | Context Window | Reasoning Modes | Multimodal Input | Primary Role / Positioning |
|---|---|---|---|---|
| Kimi K2.8 Preview | 1,048,576 tokens (All tiers) | Configurable (low, high, max) | Text, Code, Image, Video | High-throughput agentic daily driver approaching K3 quality |
| Kimi K2.7 Code | 256,000 tokens | Fixed / Limited | Text, Code | Previous-generation coding model with higher reasoning overhead |
| Kimi K3 | 1,048,576 tokens (Enterprise) | Extended thinking | Text, Code, Multimodal | 2.8T MoE flagship frontier model |
| Claude Code (Anthropic) | 200,000 tokens | Extended thinking | Text, Code, Image | Western benchmark for terminal-based agentic software engineering |
Technical Deep Dive
- Context Window Architecture: Expanded to a full 1,048,576 tokens (1M tokens) accessible across all service tiers without enterprise gating.
- Endpoint Continuity: Drop-in deployment utilizing the static
kimi-for-codingAPI endpoint alias, preserving compatibility with existing CLI tools, IDE plugins, and CI/CD pipelines. - Configurable Reasoning Engine: Exposes tiered thinking effort controls (
low,high,max), allowing programmatic balancing of inference latency and reasoning depth. - Intelligent Fallback Routing: Architectural fallback redirects any queries with disabled thinking on either K3 or K2.8 to run directly on the K2.8 non-thinking stack.
- Multimodal Ingestion: Native multi-format parsing capable of processing repositories, documentation, images, and video files in-context.
- Throughput Performance: Optimized inference serving reaching up to ~260 tokens/s on developer infrastructure.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-07Moonshot AI launches flagship Kimi K3, a 2.8T MoE frontier model with 1M context
- 2026-09Kimi K2.8 Preview first noted in official Kimi Code changelog on September 11
- 2026-09Full public rollout of Kimi K2.8 Preview across Kimi Code and Kimi Work announced
Sources (10)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.



