🐯Stalecollected in 87m

DeepSeek V4 Slashes Prices, Native Huawei Support

DeepSeek V4 Slashes Prices, Native Huawei Support
PostLinkedIn
🐯Read original on 虎嗅

💡DeepSeek V4: cheapest API + Huawei native, beats Claude on agents

⚡ 30-Second TL;DR

What Changed

V4-Flash cache input 0.02元/M tokens, Pro 0.025元 – lowest in market

Why It Matters

Undercuts global models on price amid shortages, validates国产 calc for scalable AI. Enables devs to run commercial apps cheaply on domestic infra.

What To Do Next

Test DeepSeek-V4-Pro API before May 5 for 75% discount on agent tasks.

Who should care:Developers & AI Engineers

Key Points

  • V4-Flash cache input 0.02元/M tokens, Pro 0.025元 – lowest in market
  • SOTA open model in Agentic Coding, rivals Claude Opus 4.6
  • Architecture: CSA+HCA cuts long-context compute 73%, Muon optimizer
  • First full CANN migration for国产 AI inference on Ascend chips

🧠 Deep Insight

AI-generated analysis for this event.

🔑 Enhanced Key Takeaways

  • DeepSeek's integration with Huawei Ascend 950PR marks the first time a major Chinese LLM provider has achieved full-stack optimization for the CANN (Compute Architecture for Neural Networks) ecosystem, effectively bypassing reliance on CUDA-based hardware for high-performance inference.
  • The 62% surge in API calls following the price reduction has triggered significant strain on existing domestic GPU cluster capacity, forcing DeepSeek to implement dynamic load balancing across heterogeneous hardware environments.
  • The adoption of the Muon optimizer in V4 represents a shift toward memory-efficient training and inference, specifically targeting the reduction of activation memory overhead which has historically been a bottleneck for long-context windows on domestic hardware.
📊 Competitor Analysis▸ Show
FeatureDeepSeek V4Qwen-Max (Alibaba)Yi-Lightning (01.AI)
Primary HardwareAscend 950PR / NVIDIANVIDIA H800/A800NVIDIA H800
Input Pricing (per M tokens)0.02元 (Flash)~0.04元~0.035元
Agentic Coding SOTAYes (Rivals Opus 4.6)HighModerate
CANN Native SupportYesPartialNo

🛠️ Technical Deep Dive

  • CSA (Context-Sparse Attention): A novel attention mechanism that dynamically prunes non-essential tokens during the KV-cache generation phase, contributing to the 73% compute reduction.
  • HCA (Hierarchical Context Aggregation): A multi-level compression technique that aggregates long-range dependencies into compact latent representations before final decoding.
  • CANN Migration: Implementation utilizes custom operator fusion kernels specifically written for the Ascend 950PR's NPU architecture, bypassing standard PyTorch-to-Ascend translation layers for lower latency.
  • Muon Optimizer: A second-order optimization technique adapted for distributed training that reduces the need for large-scale optimizer state storage, allowing for larger effective batch sizes on memory-constrained hardware.

🔮 Future ImplicationsAI analysis grounded in cited sources

DeepSeek will achieve parity with NVIDIA-based inference costs by Q4 2026.
The successful optimization of CANN-native kernels allows DeepSeek to leverage cheaper, domestically available NPU supply chains without sacrificing throughput.
Major Chinese cloud providers will be forced to lower API pricing by at least 30% within 60 days.
DeepSeek's aggressive pricing strategy has established a new market floor that competitors cannot ignore without risking significant developer churn.

Timeline

2024-01
DeepSeek releases its first open-source model series, establishing its focus on high-efficiency architectures.
2025-05
DeepSeek V3 launch, introducing initial support for domestic hardware acceleration.
2026-02
DeepSeek announces strategic partnership with Huawei to optimize model inference on Ascend NPUs.
2026-04
DeepSeek V4 release with full CANN migration and aggressive API price cuts.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅