🔥Stalecollected in 27m

DeepSeek-V4-Pro API announces permanent price reduction

DeepSeek-V4-Pro API announces permanent price reduction
PostLinkedIn
🔥Read original on 36氪

💡Significant API price reduction for high-performance LLMs.

⚡ 30-Second TL;DR

What Changed

Permanent price cut to 25% of original cost

Why It Matters

This aggressive pricing strategy significantly lowers the barrier for developers to integrate high-performance models into production environments.

What To Do Next

Update your long-term cloud infrastructure budget and cost-per-token projections based on the new DeepSeek-V4-Pro pricing.

Who should care:Developers & AI Engineers

Key Points

  • Permanent price cut to 25% of original cost
  • Effective after May 31, 2026
  • Follows the end of the 2.5x discount promotion

🧠 Deep Insight

Web-grounded analysis with 16 cited sources.

🔑 Enhanced Key Takeaways

  • The permanent price adjustment for DeepSeek-V4-Pro will set its cost at approximately $0.435 per million input tokens and $0.87 per million output tokens, a 75% reduction from its original pricing of $1.74 per million input and $3.48 per million output tokens.
  • In addition to the V4-Pro price cut, DeepSeek has also permanently reduced the cost of input cache hits across its entire API portfolio to one-tenth of their previous levels, effective immediately.
  • This aggressive pricing strategy is aimed at intensifying competition within the global AI market, directly challenging established US AI providers like OpenAI, Google, and Anthropic by offering significantly lower per-token costs.
  • DeepSeek-V4-Pro is a Mixture-of-Experts (MoE) model featuring 1.6 trillion total parameters with 49 billion active parameters, and supports an extensive 1 million token context window.
  • The model incorporates a hybrid attention architecture (Compressed Sparse Attention and Heavily Compressed Attention) and Manifold-Constrained Hyper-Connections (mHC) to enhance long-context efficiency, requiring only 27% of single-token inference FLOPs and 10% of KV cache compared to DeepSeek-V3.2.
📊 Competitor Analysis▸ Show
ModelInput Price (per 1M tokens)Output Price (per 1M tokens)Context WindowKey Benchmarks (SWE-bench Verified)
DeepSeek-V4-Pro (Permanent Price)$0.435$0.871M tokens80.6%
DeepSeek-V4-Pro (Original Price)$1.74$3.481M tokens80.6%
DeepSeek-V4-Flash$0.14$0.281M tokens79.0%
DeepSeek V3.2 (legacy)$0.27$1.1 / $0.42128K tokens~69%
OpenAI GPT-5.5$5.00$30.001M tokensN/A
Google Gemini 3.1 Pro$2.00$12.001M tokensN/A
Anthropic Claude Opus 4.7$5.00$25.001M tokens80.8% (Claude Opus 4.6)

🛠️ Technical Deep Dive

  • DeepSeek-V4-Pro is a Mixture-of-Experts (MoE) language model with 1.6 trillion total parameters and 49 billion activated parameters.
  • It features a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to significantly improve long-context efficiency.
  • This architecture results in DeepSeek-V4-Pro requiring only 27% of single-token inference FLOPs and 10% of KV cache compared to DeepSeek-V3.2 when processing a 1 million token context.
  • The model also incorporates Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across its deep layer stack while maintaining expressivity.
  • DeepSeek-V4-Pro supports a maximum context length of 1 million tokens and can generate up to 384,000 tokens per response.
  • Its post-training pipeline involves two stages: independent cultivation of domain-specific experts (using SFT and GRPO) followed by unified model consolidation through on-policy distillation.
  • The model offers three reasoning modes: Non-think (for fast processing), Think High (for logical analysis), and Think Max (for full reasoning extent).
  • DeepSeek-V4-Pro has been optimized to run on Huawei's Ascend AI processors, indicating a strategic shift from its prior reliance on Nvidia hardware.

🔮 Future ImplicationsAI analysis grounded in cited sources

The permanent price reduction will intensify the ongoing AI price war.
DeepSeek's aggressive pricing strategy, particularly for a high-performing model like V4-Pro, will likely pressure competitors to re-evaluate and potentially lower their own API costs to remain competitive in the market.
DeepSeek will gain significant market share among developers and enterprises.
The combination of near-frontier performance, open-weight availability, a large context window, and substantially reduced pricing makes DeepSeek-V4-Pro a highly attractive and cost-effective option for a wide range of AI applications and users.
The move will accelerate the adoption of open-weight models in production environments.
By making a powerful model like V4-Pro more affordable and accessible, DeepSeek lowers the barrier to entry for startups and smaller enterprises, fostering innovation and broader deployment of advanced AI.

Timeline

2016-02
High-Flyer, DeepSeek's parent hedge fund, co-founded by Liang Wenfeng.
2023-07
DeepSeek Artificial Intelligence founded by Liang Wenfeng.
2023-11
DeepSeek Coder, an open-source model for coding tasks, released.
2024-05
DeepSeek-V2 chatbot model released, gaining popularity in China for cost-efficiency.
2025-01
DeepSeek-R1 reasoning model and eponymous chatbot application launched, achieving international prominence.
2026-04
DeepSeek-V4-Pro and DeepSeek-V4-Flash models released.

📎 Sources (16)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. qz.com
  2. kilo.ai
  3. thenextweb.com
  4. requesty.ai
  5. inworld.ai
  6. mexc.com
  7. phemex.com
  8. mashable.com
  9. deepinfra.com
  10. openrouter.ai
  11. nvidia.com
  12. huggingface.co
  13. lightning.ai
  14. datacamp.com
  15. devtk.ai
  16. pricepertoken.com
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪