DeepSeek Implements Permanent Price Cuts

💡DeepSeek's permanent price cut is reshaping the unit economics of AI. See if your infrastructure costs can be optimized.
⚡ 30-Second TL;DR
What Changed
DeepSeek officially shifts to a permanent low-price model
Why It Matters
This aggressive pricing forces other LLM providers to re-evaluate their unit economics and potentially triggers a broader price war in the AI infrastructure sector.
What To Do Next
Re-calculate your projected inference costs using DeepSeek's new pricing to determine if migrating your production workloads is financially viable.
Key Points
- •DeepSeek officially shifts to a permanent low-price model
- •Analysts link the pricing strategy to a $10 trillion market opportunity
- •Amazon executives are analyzing the cost-efficiency behind DeepSeek's pricing
🧠 Deep Insight
Web-grounded analysis with 14 cited sources.
🔑 Enhanced Key Takeaways
- •DeepSeek's permanent price reduction for its V4 Pro model is a 75% cut, making a previous promotional offer permanent that was originally set to expire on May 31, 2026.
- •The new permanent pricing for DeepSeek V4 Pro is set at approximately $0.003625 per million input tokens and $0.87 per million output tokens, positioning it to significantly undercut competitors.
- •This aggressive pricing strategy is partly enabled by DeepSeek's optimization for Huawei's Ascend 950 AI chips, reducing its reliance on NVIDIA hardware affected by U.S. export restrictions.
- •DeepSeek's move is seen as a direct challenge to Western AI firms like OpenAI and Google, aiming to capture market share by offering a more affordable alternative for enterprise and power users who consume millions of tokens daily.
- •The company's decision to lock in the discount one month after launching the V4 models suggests a strategic prioritization of market share over immediate per-unit revenue, emphasizing a 'cost-effective 1M context length' era.
📊 Competitor Analysis▸ Show
| Provider | Model | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|---|
| DeepSeek | V4 Pro (New Permanent Pricing) | $0.003625 | $0.87 |
| DeepSeek | V3.2 | $0.28 | $0.42 |
| OpenAI | GPT-5 | $2.50 | $10.00 |
| Anthropic | Claude Opus 4.7 | $5.00 | $25.00 |
| Gemini 3.5 Flash | $0.15 | $0.60 |
🛠️ Technical Deep Dive
- DeepSeek-V3 and DeepSeek-R1 models utilize a Mixture-of-Experts (MoE) architecture, comprising 671 billion total parameters with only 37 billion activated per token, which enhances efficiency during inference and training.
- Key architectural innovations include Multi-head Latent Attention (MLA), designed to optimize attention operations and reduce memory consumption by compressing key-value pairs into a low-dimensional latent space.
- DeepSeek-V3.2 further incorporates DeepSeek Sparse Attention (DSA), which achieves a 70% reduction in computational complexity compared to standard attention mechanisms through learned sparsity patterns.
- The models employ FP8 mixed precision for training, which doubles compute efficiency and halves memory usage compared to BF16, contributing to lower training costs.
- The DeepSeekMoE component, specifically in V3, uses 256 routed experts and 1 shared expert, with each token dynamically interacting with 8 specialized experts plus the single shared expert, alongside an auxiliary-loss-free load balancing strategy.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (14)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 钛媒体 ↗
