DeepSeek API Cache Pricing Drops to 1/10th

💡DeepSeek cache input now global cheapest: 0.025¥/M tokens, 90% off launch!
⚡ 30-Second TL;DR
What Changed
Input cache hit price slashed to 1/10 of original launch price.
Why It Matters
Positions DeepSeek as top cost leader in LLM APIs, boosting adoption for high-volume, cache-reliant apps. Heightens price competition, pressuring rivals like global providers to match. Benefits devs scaling inference economically.
What To Do Next
Test DeepSeek V4-Pro API cache endpoints for your workloads to achieve 90% input savings.
Key Points
- •Input cache hit price slashed to 1/10 of original launch price.
- •V4-Pro cached input: 0.025 CNY per million tokens with discounts.
- •Covers full DeepSeek-V4-Pro and V4-Flash model series.
- •Sets new global lowest price for LLM API caching.
🧠 Deep Insight
AI-generated analysis for this event — not the original article.
🔑 Enhanced Key Takeaways
- •The price reduction is part of DeepSeek's broader strategy to incentivize the adoption of Context Caching, a feature designed to reduce latency and costs for applications with long, repetitive prompts or large system instructions.
- •This aggressive pricing model is specifically optimized for DeepSeek's Mixture-of-Experts (MoE) architecture, which allows for efficient token processing when cached segments are reused across multiple API requests.
- •Industry analysts suggest this move is a direct response to increasing commoditization in the LLM market, aiming to capture high-volume enterprise developers who prioritize cost-efficiency for RAG (Retrieval-Augmented Generation) pipelines.
📊 Competitor Analysis▸ Show
| Feature | DeepSeek V4-Pro (Cached) | OpenAI GPT-4o (Cached) | Anthropic Claude 3.5 Sonnet (Cached) |
|---|---|---|---|
| Pricing (per 1M tokens) | 0.025 CNY (~$0.0035 USD) | ~$1.25 USD | ~$1.50 USD |
| Architecture | Mixture-of-Experts (MoE) | Dense/MoE Hybrid | Dense |
| Primary Advantage | Lowest cost for high-volume caching | Ecosystem integration | Context window performance |
🛠️ Technical Deep Dive
- •Context Caching implementation: DeepSeek utilizes a KV (Key-Value) cache storage mechanism that persists across API calls, allowing the model to skip the computation of pre-processed prompt prefixes.
- •MoE Efficiency: The V4 series employs a sparse activation mechanism where only a subset of parameters is active per token, significantly reducing the computational overhead when the cache hit rate is high.
- •Integration: The API requires developers to explicitly define a 'cache_id' or 'cache_config' in the request header to trigger the hit, ensuring granular control over which prompt segments are stored.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: cnBeta (Full RSS) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.