DeepSeek Makes 75% Price Cut Permanent, Escalating AI War

๐กDeepSeek's permanent 75% price cut is a major market shift that could significantly lower your AI infrastructure costs.
โก 30-Second TL;DR
What Changed
DeepSeek V4 Pro pricing is now permanently reduced by 75%.
Why It Matters
This move forces other AI providers to reconsider their pricing models to remain competitive. Developers and startups can expect lower operational costs for inference-heavy applications.
What To Do Next
Benchmark your current LLM inference costs against DeepSeek's new pricing to determine if migrating specific workloads could optimize your cloud spend.
Key Points
- โขDeepSeek V4 Pro pricing is now permanently reduced by 75%.
- โขNew token pricing ranges from $0.003625 to $0.87 per million tokens.
- โขThe move directly challenges major providers like OpenAI by significantly lowering the barrier to entry for high-performance models.
๐ง Deep Insight
Web-grounded analysis with 21 cited sources.
๐ Enhanced Key Takeaways
- โขThe permanent price cut positions DeepSeek V4 Pro's output tokens at $0.87 per million, making it approximately eight to nine times cheaper than OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 for comparable output token costs.
- โขDeepSeek V4 Pro is an open-weight Mixture-of-Experts (MoE) model featuring a 1 million token context window and the ability to output up to 384,000 tokens in a single request, matching or exceeding the specifications of several Western competitors.
- โขThis aggressive pricing strategy is consistent with DeepSeek's previous market entries, as the company similarly made a substantial promotional discount for its DeepSeek V3 model permanent at the end of 2024.
- โขDeepSeek V4 Pro demonstrates strong performance on agentic benchmarks, scoring around 91.2% on SWE-Bench Verified and 93.5% on LiveCodeBench, placing it in a similar performance tier to Claude Opus 4.7 and GPT-5.5 for complex coding and multi-step reasoning tasks.
๐ Competitor Analysisโธ Show
| Feature/Model | DeepSeek V4 Pro | OpenAI GPT-5.5 | Anthropic Claude Opus 4.7 | Google Gemini 3.1 Pro |
|---|---|---|---|---|
| Pricing (per 1M tokens) | ||||
| Input | $0.435 | $5.00 | $5.00 | $2.00 |
| Output | $0.87 | $30.00 | $25.00 | $12.00 |
| Context Window | 1M tokens | 1M tokens | 1M tokens | 2M tokens |
| Key Benchmarks | ||||
| SWE-Bench Verified | ~91.2% | Slightly above DeepSeek V4 Pro | ~93.9% | N/A |
| LiveCodeBench | 93.5% | N/A | 88.8% | N/A |
| CodeForces | #1 | N/A | N/A | N/A |
| Architecture | MoE (1.6T total, 49B active) | Proprietary | Proprietary | Proprietary |
| Open Weights | Yes (MIT License) | No | No | No |
๐ ๏ธ Technical Deep Dive
- Model Architecture: DeepSeek V4 Pro is a Mixture-of-Experts (MoE) model with 1.6 trillion total parameters and 49 billion activated parameters.
- Attention Mechanism: It incorporates a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to enhance long-context efficiency. This design reduces single-token inference FLOPs by 27% and KV cache by 10% compared to DeepSeek-V3.2 at a 1M-token context.
- Connections & Optimization: The model utilizes Manifold-Constrained Hyper-Connections (mHC) to improve signal propagation stability across layers and employs the Muon optimizer for faster convergence and greater training stability.
- Post-training Pipeline: DeepSeek V4 Pro's post-training involves a two-stage pipeline: independent domain-expert cultivation (using SFT + GRPO) followed by unified model consolidation via on-policy distillation.
- Reasoning Modes: It supports three reasoning modes: Non-think (fast), Think High (logical analysis), and Think Max (full reasoning extent), allowing for flexible computational intensity based on task requirements.
- Output Capabilities: The model supports structured JSON output, as well as function and tool calling.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (21)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) โ


