🔥36氪•Stalecollected in 27m
DeepSeek-V4-Pro API announces permanent price reduction
💡Significant API price reduction for high-performance LLMs.
⚡ 30-Second TL;DR
What Changed
Permanent price cut to 25% of original cost
Why It Matters
This aggressive pricing strategy significantly lowers the barrier for developers to integrate high-performance models into production environments.
What To Do Next
Update your long-term cloud infrastructure budget and cost-per-token projections based on the new DeepSeek-V4-Pro pricing.
Who should care:Developers & AI Engineers
Key Points
- •Permanent price cut to 25% of original cost
- •Effective after May 31, 2026
- •Follows the end of the 2.5x discount promotion
🧠 Deep Insight
Web-grounded analysis with 16 cited sources.
🔑 Enhanced Key Takeaways
- •The permanent price adjustment for DeepSeek-V4-Pro will set its cost at approximately $0.435 per million input tokens and $0.87 per million output tokens, a 75% reduction from its original pricing of $1.74 per million input and $3.48 per million output tokens.
- •In addition to the V4-Pro price cut, DeepSeek has also permanently reduced the cost of input cache hits across its entire API portfolio to one-tenth of their previous levels, effective immediately.
- •This aggressive pricing strategy is aimed at intensifying competition within the global AI market, directly challenging established US AI providers like OpenAI, Google, and Anthropic by offering significantly lower per-token costs.
- •DeepSeek-V4-Pro is a Mixture-of-Experts (MoE) model featuring 1.6 trillion total parameters with 49 billion active parameters, and supports an extensive 1 million token context window.
- •The model incorporates a hybrid attention architecture (Compressed Sparse Attention and Heavily Compressed Attention) and Manifold-Constrained Hyper-Connections (mHC) to enhance long-context efficiency, requiring only 27% of single-token inference FLOPs and 10% of KV cache compared to DeepSeek-V3.2.
📊 Competitor Analysis▸ Show
| Model | Input Price (per 1M tokens) | Output Price (per 1M tokens) | Context Window | Key Benchmarks (SWE-bench Verified) |
|---|---|---|---|---|
| DeepSeek-V4-Pro (Permanent Price) | $0.435 | $0.87 | 1M tokens | 80.6% |
| DeepSeek-V4-Pro (Original Price) | $1.74 | $3.48 | 1M tokens | 80.6% |
| DeepSeek-V4-Flash | $0.14 | $0.28 | 1M tokens | 79.0% |
| DeepSeek V3.2 (legacy) | $0.27 | $1.1 / $0.42 | 128K tokens | ~69% |
| OpenAI GPT-5.5 | $5.00 | $30.00 | 1M tokens | N/A |
| Google Gemini 3.1 Pro | $2.00 | $12.00 | 1M tokens | N/A |
| Anthropic Claude Opus 4.7 | $5.00 | $25.00 | 1M tokens | 80.8% (Claude Opus 4.6) |
🛠️ Technical Deep Dive
- DeepSeek-V4-Pro is a Mixture-of-Experts (MoE) language model with 1.6 trillion total parameters and 49 billion activated parameters.
- It features a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to significantly improve long-context efficiency.
- This architecture results in DeepSeek-V4-Pro requiring only 27% of single-token inference FLOPs and 10% of KV cache compared to DeepSeek-V3.2 when processing a 1 million token context.
- The model also incorporates Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across its deep layer stack while maintaining expressivity.
- DeepSeek-V4-Pro supports a maximum context length of 1 million tokens and can generate up to 384,000 tokens per response.
- Its post-training pipeline involves two stages: independent cultivation of domain-specific experts (using SFT and GRPO) followed by unified model consolidation through on-policy distillation.
- The model offers three reasoning modes: Non-think (for fast processing), Think High (for logical analysis), and Think Max (for full reasoning extent).
- DeepSeek-V4-Pro has been optimized to run on Huawei's Ascend AI processors, indicating a strategic shift from its prior reliance on Nvidia hardware.
🔮 Future ImplicationsAI analysis grounded in cited sources
The permanent price reduction will intensify the ongoing AI price war.
DeepSeek's aggressive pricing strategy, particularly for a high-performing model like V4-Pro, will likely pressure competitors to re-evaluate and potentially lower their own API costs to remain competitive in the market.
DeepSeek will gain significant market share among developers and enterprises.
The combination of near-frontier performance, open-weight availability, a large context window, and substantially reduced pricing makes DeepSeek-V4-Pro a highly attractive and cost-effective option for a wide range of AI applications and users.
The move will accelerate the adoption of open-weight models in production environments.
By making a powerful model like V4-Pro more affordable and accessible, DeepSeek lowers the barrier to entry for startups and smaller enterprises, fostering innovation and broader deployment of advanced AI.
⏳ Timeline
2016-02
High-Flyer, DeepSeek's parent hedge fund, co-founded by Liang Wenfeng.
2023-07
DeepSeek Artificial Intelligence founded by Liang Wenfeng.
2023-11
DeepSeek Coder, an open-source model for coding tasks, released.
2024-05
DeepSeek-V2 chatbot model released, gaining popularity in China for cost-efficiency.
2025-01
DeepSeek-R1 reasoning model and eponymous chatbot application launched, achieving international prominence.
2026-04
DeepSeek-V4-Pro and DeepSeek-V4-Flash models released.
📎 Sources (16)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗