DeepSeek V4 Slashes Prices, Native Huawei Support

💡DeepSeek V4: cheapest API + Huawei native, beats Claude on agents
⚡ 30-Second TL;DR
What Changed
V4-Flash cache input 0.02元/M tokens, Pro 0.025元 – lowest in market
Why It Matters
Undercuts global models on price amid shortages, validates国产 calc for scalable AI. Enables devs to run commercial apps cheaply on domestic infra.
What To Do Next
Test DeepSeek-V4-Pro API before May 5 for 75% discount on agent tasks.
Key Points
- •V4-Flash cache input 0.02元/M tokens, Pro 0.025元 – lowest in market
- •SOTA open model in Agentic Coding, rivals Claude Opus 4.6
- •Architecture: CSA+HCA cuts long-context compute 73%, Muon optimizer
- •First full CANN migration for国产 AI inference on Ascend chips
🧠 Deep Insight
AI-generated analysis for this event — not the original article.
🔑 Enhanced Key Takeaways
- •DeepSeek's integration with Huawei Ascend 950PR marks the first time a major Chinese LLM provider has achieved full-stack optimization for the CANN (Compute Architecture for Neural Networks) ecosystem, effectively bypassing reliance on CUDA-based hardware for high-performance inference.
- •The 62% surge in API calls following the price reduction has triggered significant strain on existing domestic GPU cluster capacity, forcing DeepSeek to implement dynamic load balancing across heterogeneous hardware environments.
- •The adoption of the Muon optimizer in V4 represents a shift toward memory-efficient training and inference, specifically targeting the reduction of activation memory overhead which has historically been a bottleneck for long-context windows on domestic hardware.
📊 Competitor Analysis▸ Show
| Feature | DeepSeek V4 | Qwen-Max (Alibaba) | Yi-Lightning (01.AI) |
|---|---|---|---|
| Primary Hardware | Ascend 950PR / NVIDIA | NVIDIA H800/A800 | NVIDIA H800 |
| Input Pricing (per M tokens) | 0.02元 (Flash) | ~0.04元 | ~0.035元 |
| Agentic Coding SOTA | Yes (Rivals Opus 4.6) | High | Moderate |
| CANN Native Support | Yes | Partial | No |
🛠️ Technical Deep Dive
- •CSA (Context-Sparse Attention): A novel attention mechanism that dynamically prunes non-essential tokens during the KV-cache generation phase, contributing to the 73% compute reduction.
- •HCA (Hierarchical Context Aggregation): A multi-level compression technique that aggregates long-range dependencies into compact latent representations before final decoding.
- •CANN Migration: Implementation utilizes custom operator fusion kernels specifically written for the Ascend 950PR's NPU architecture, bypassing standard PyTorch-to-Ascend translation layers for lower latency.
- •Muon Optimizer: A second-order optimization technique adapted for distributed training that reduces the need for large-scale optimizer state storage, allowing for larger effective batch sizes on memory-constrained hardware.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.



