DeepSeek Makes 75% Discount on V4-Pro Permanent
๐กDeepSeek's permanent price slash significantly alters the cost-benefit analysis for scaling LLM-based applications.
โก 30-Second TL;DR
What Changed
DeepSeek V4-Pro pricing is now permanently reduced by 75%.
Why It Matters
This aggressive pricing strategy forces other model providers to re-evaluate their API costs to remain competitive. It significantly lowers the barrier to entry for developers building high-scale applications on DeepSeek's infrastructure.
What To Do Next
Update your cloud infrastructure budget and API cost projections to reflect the permanent 75% reduction in DeepSeek V4-Pro usage costs.
Key Points
- โขDeepSeek V4-Pro pricing is now permanently reduced by 75%.
- โขDeveloper costs remain fixed at 25% of the original launch price.
- โขThe decision signals a long-term commitment to aggressive AI model pricing competition.
๐ง Deep Insight
Web-grounded analysis with 14 cited sources.
๐ Enhanced Key Takeaways
- โขThe permanent 75% discount for DeepSeek V4-Pro was a strategic move, initially a promotional campaign in April 2026, made permanent in response to widespread developer frustration over restrictive usage caps imposed by Western AI chatbots like Google Gemini, Anthropic's Claude, and Perplexity.
- โขIn addition to the V4-Pro price cut, DeepSeek also implemented a permanent 90% cost reduction for input cache hits across its entire API lineup, significantly lowering expenses for repetitive prompts and continuous system instructions.
- โขWith the permanent discount, DeepSeek V4-Pro's non-cached input tokens are priced at $0.435 per million tokens (down from an original $1.74), and output tokens are set at $0.87 per million tokens (down from $3.48).
- โขDeepSeek V4-Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model with 49 billion active parameters, designed for advanced reasoning, complex software engineering, and long-running agentic tasks, and supports a 1 million token context window.
๐ Competitor Analysisโธ Show
| Feature/Model | DeepSeek V4-Pro (Discounted) | Anthropic Claude Opus 4.7 | OpenAI GPT-4o | Mistral Large 2 |
|---|---|---|---|---|
| Input Price (per 1M tokens) | $0.435 (non-cached) | $5.00 | $2.50 | $3.00 |
| Output Price (per 1M tokens) | $0.87 | $25.00 | $10.00 | $9.00 |
| Context Window | 1M tokens | 1M tokens | 128K tokens | 128K tokens |
| Total Parameters | 1.6T (49B active MoE) | N/A (Closed-source) | N/A (Closed-source) | N/A (Closed-source) |
| Open Weights | Yes (MIT License) | No | No | No |
| SWE-bench Verified | 80.6% | 64.3% (SWE-bench Pro) | N/A | N/A |
| LiveCodeBench | 93.5% | 88.8% | N/A | N/A |
| Codeforces Rating | 3206 | N/A | N/A | N/A |
๐ ๏ธ Technical Deep Dive
- DeepSeek V4-Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model with 49 billion active parameters.
- It features a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to enhance long-context efficiency.
- This architecture reduces single-token inference FLOPs by 73% and KV cache usage by 90% compared to DeepSeek-V3.2 at a 1M-token context.
- The model incorporates Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across its deep layer stack while preserving expressivity.
- DeepSeek V4-Pro utilizes the Muon optimizer, contributing to faster convergence and improved training stability.
- It was pre-trained on more than 32 trillion diverse and high-quality tokens.
- The post-training process involves a two-stage pipeline: initial cultivation of independent domain-specific experts (using SFT + GRPO) followed by unified model consolidation via on-policy distillation.
- The model supports a 1 million token context window and offers three reasoning modes: Non-think (fast), Think High (logical analysis), and Think Max (full reasoning extent).
- DeepSeek V4-Pro is compatible with the Hugging Face Transformers library and vLLM for efficient inference.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (14)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Bloomberg Technology โ