๐Ÿ“ŠStalecollected in 32m

DeepSeek Makes 75% Discount on V4-Pro Permanent

PostLinkedIn
๐Ÿ“ŠRead original on Bloomberg Technology

๐Ÿ’กDeepSeek's permanent price slash significantly alters the cost-benefit analysis for scaling LLM-based applications.

โšก 30-Second TL;DR

What Changed

DeepSeek V4-Pro pricing is now permanently reduced by 75%.

Why It Matters

This aggressive pricing strategy forces other model providers to re-evaluate their API costs to remain competitive. It significantly lowers the barrier to entry for developers building high-scale applications on DeepSeek's infrastructure.

What To Do Next

Update your cloud infrastructure budget and API cost projections to reflect the permanent 75% reduction in DeepSeek V4-Pro usage costs.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขDeepSeek V4-Pro pricing is now permanently reduced by 75%.
  • โ€ขDeveloper costs remain fixed at 25% of the original launch price.
  • โ€ขThe decision signals a long-term commitment to aggressive AI model pricing competition.

๐Ÿง  Deep Insight

Web-grounded analysis with 14 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขThe permanent 75% discount for DeepSeek V4-Pro was a strategic move, initially a promotional campaign in April 2026, made permanent in response to widespread developer frustration over restrictive usage caps imposed by Western AI chatbots like Google Gemini, Anthropic's Claude, and Perplexity.
  • โ€ขIn addition to the V4-Pro price cut, DeepSeek also implemented a permanent 90% cost reduction for input cache hits across its entire API lineup, significantly lowering expenses for repetitive prompts and continuous system instructions.
  • โ€ขWith the permanent discount, DeepSeek V4-Pro's non-cached input tokens are priced at $0.435 per million tokens (down from an original $1.74), and output tokens are set at $0.87 per million tokens (down from $3.48).
  • โ€ขDeepSeek V4-Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model with 49 billion active parameters, designed for advanced reasoning, complex software engineering, and long-running agentic tasks, and supports a 1 million token context window.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature/ModelDeepSeek V4-Pro (Discounted)Anthropic Claude Opus 4.7OpenAI GPT-4oMistral Large 2
Input Price (per 1M tokens)$0.435 (non-cached)$5.00$2.50$3.00
Output Price (per 1M tokens)$0.87$25.00$10.00$9.00
Context Window1M tokens1M tokens128K tokens128K tokens
Total Parameters1.6T (49B active MoE)N/A (Closed-source)N/A (Closed-source)N/A (Closed-source)
Open WeightsYes (MIT License)NoNoNo
SWE-bench Verified80.6%64.3% (SWE-bench Pro)N/AN/A
LiveCodeBench93.5%88.8%N/AN/A
Codeforces Rating3206N/AN/AN/A

๐Ÿ› ๏ธ Technical Deep Dive

  • DeepSeek V4-Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model with 49 billion active parameters.
  • It features a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to enhance long-context efficiency.
  • This architecture reduces single-token inference FLOPs by 73% and KV cache usage by 90% compared to DeepSeek-V3.2 at a 1M-token context.
  • The model incorporates Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across its deep layer stack while preserving expressivity.
  • DeepSeek V4-Pro utilizes the Muon optimizer, contributing to faster convergence and improved training stability.
  • It was pre-trained on more than 32 trillion diverse and high-quality tokens.
  • The post-training process involves a two-stage pipeline: initial cultivation of independent domain-specific experts (using SFT + GRPO) followed by unified model consolidation via on-policy distillation.
  • The model supports a 1 million token context window and offers three reasoning modes: Non-think (fast), Think High (logical analysis), and Think Max (full reasoning extent).
  • DeepSeek V4-Pro is compatible with the Hugging Face Transformers library and vLLM for efficient inference.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

DeepSeek's aggressive pricing will intensify the global AI price war, particularly impacting Western providers.
By permanently offering significantly lower costs, DeepSeek pressures competitors to reduce their own pricing to remain competitive for developers and enterprises.
The focus on cost-efficiency and open weights will accelerate AI adoption in agentic workflows and large-scale code analysis.
Lower costs for long context windows and cached inputs make complex, iterative AI applications economically viable for a broader range of developers and businesses.
DeepSeek will gain significant market share among developers and enterprises dissatisfied with Western AI models' usage restrictions.
The permanent discount and higher usage limits directly address a key pain point for users currently facing restrictive caps from competitors.

โณ Timeline

2023-05
DeepSeek AI founded by Liang Wenfeng, backed by High-Flyer.
2023-11
DeepSeek released its first model, DeepSeek Coder.
2024-05
DeepSeek-V2 launched, gaining popularity and triggering a price war in China.
2025-01
DeepSeek launched its DeepSeek-R1 model and a mobile chatbot, becoming the most downloaded app on the U.S. iOS App Store.
2026-04
DeepSeek released the V4 series (V4-Pro and V4-Flash) with an initial 75% promotional discount on V4-Pro and a 90% reduction on input cache hits.
2026-05
DeepSeek announced the 75% price reduction for its V4-Pro model would be permanent.

๐Ÿ“Ž Sources (14)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. androidheadlines.com
  2. thenextweb.com
  3. devtk.ai
  4. deepinfra.com
  5. nvidia.com
  6. huggingface.co
  7. datacamp.com
  8. pricepertoken.com
  9. docsbot.ai
  10. docsbot.ai
  11. docsbot.ai
  12. lightning.ai
  13. huggingface.co
  14. scmp.com
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Bloomberg Technology โ†—