Huawei chips successfully train DeepSeek-V4-Pro model

๐กHuawei chips are now training complex models, signaling a shift in the global AI hardware supply chain.
โก 30-Second TL;DR
What Changed
Huawei Ascend 910C chips were used for post-training the DeepSeek-V4-Pro model.
Why It Matters
This development suggests that domestic Chinese hardware is becoming a viable alternative for training large-scale models, potentially altering the competitive landscape for AI infrastructure in the region.
What To Do Next
Monitor the performance benchmarks of Ascend 910C against Nvidia H100 to assess if your infrastructure strategy needs to account for non-US hardware alternatives.
Key Points
- โขHuawei Ascend 910C chips were used for post-training the DeepSeek-V4-Pro model.
- โขThe milestone demonstrates progress in China's ability to perform complex AI training domestically.
- โขThis reduces reliance on high-end US-made GPUs for advanced AI development.
๐ง Deep Insight
Web-grounded analysis with 24 cited sources.
๐ Enhanced Key Takeaways
- โขThe DeepSeek-V4-Pro model is a Mixture-of-Experts (MoE) architecture featuring 1.6 trillion total parameters and 49 billion activated parameters, designed to support a substantial 1 million token context window.
- โขHuawei's Ascend 910C chip utilizes a dual-die packaging design, integrating two Ascend 910B SoCs, and is manufactured using SMIC's 7nm (N+2) process technology.
- โขDeepSeek's AI models offer native support for Huawei's Ascend processors, facilitating a straightforward conversion from CUDA to CUNN, which is vital for seamless integration of Huawei's hardware into AI development workflows.
- โขThe United States has recently intensified its export controls, issuing new guidance to prevent Chinese companies from circumventing restrictions by accessing advanced American AI chips, such as Nvidia's Blackwell and Rubin, and AMD's MI350X, through overseas subsidiaries.
- โขChina has officially recognized AI training and inference chips, including Huawei's Ascend 910, within its 'secure and reliable' technology assessment framework, underscoring a state-backed initiative to promote domestic alternatives.
๐ Competitor Analysisโธ Show
| Feature/Metric | Huawei Ascend 910C | Nvidia H100 (Hopper) | Nvidia A100 (Ampere) |
|---|---|---|---|
| FP16 Performance | 800 TFLOPS | ~1000 TFLOPS (implied from 80% comparison) | 320 TFLOPS |
| Inference Performance | ~60% of H100 | Higher than 910C | More suitable for LLMs beyond 10B parameters |
| Memory Bandwidth | 3.2 TB/s | Higher than 910C (e.g., HBM3/HBM3e) | Higher than 910C |
| Process Technology | SMIC 7nm (N+2) | TSMC 4N (custom 5nm class) | TSMC 7nm |
| Logic Die Area | ~60% larger than H100 for 80% performance | Smaller than 910C for higher performance | Efficient design |
| TDP | 550 W | Prioritizes raw power, higher energy use | |
| Interconnect | HCCL, PCIe 4.0, 512 GB/s mesh | NVLink, high die-to-die bandwidth | NVLink |
| Software Ecosystem | MindSpore (growing) | CUDA, TensorRT (well-established) | CUDA, TensorRT (well-established) |
| Training Reliability | Critical weakness for long-term training | Undisputed lead | Strong |
| Estimated Price | ~$28,000 (similar to H100) | ~$28,000 |
๐ ๏ธ Technical Deep Dive
- DeepSeek-V4-Pro Model:
- Architecture: Mixture-of-Experts (MoE) with a novel Hybrid Attention mechanism, combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA).
- Parameters: 1.6 trillion total parameters with 49 billion activated parameters per token.
- Context Window: Supports a maximum context length of 1 million tokens.
- Training Data: Pre-trained on over 32 trillion diverse and high-quality tokens, with a focus on long documents and agentic execution traces.
- Post-training: Employs a two-stage pipeline: initial independent cultivation of domain-specific experts (via SFT and RL with GRPO), followed by unified model consolidation through on-policy distillation.
- Optimizer: Utilizes the Muon optimizer for enhanced convergence speed and training stability.
- Efficiency: Achieves significant efficiency gains, requiring only 27% of single-token inference FLOPs and 10% of KV cache compared to DeepSeek-V3.2 at a 1M-token context.
- Precision: MoE expert parameters are in FP4 precision, while most other parameters use FP8.
- Huawei Ascend 910C Chips:
- Architecture: Part of the DaVinci family, integrating multiple AI cores, each comprising high-throughput matrix-multiply cube units ('AIC'), wide vector SIMD engines ('AIV'), scalar units, and multi-level scratchpad memories.
- Process Technology: Manufactured using SMIC's 7nm (N+2) process.
- Transistor Count: Features approximately 53 billion transistors.
- Packaging: Employs a dual-die packaging design, effectively co-packaging two Ascend 910B SoCs, interconnected via an organic substrate.
- Peak Throughput: Delivers 800 TFLOPS (BF16/FP16), 100 TFLOPS (FP32), and 800 TOPS (INT8) for tensor operations.
- Memory Bandwidth: Provides 3.2 TB/s memory bandwidth.
- Interconnects: Features a 512 GB/s bidirectional mesh network across AI cores, PCIe 4.0 ร16 (64 GB/s) for host-NPU communication, and Distributed HCCL (Huawei Collective Communication Library) over PCIe or TCP for multi-node clusters; Huawei Collective Communication System (HCCS) is designed to rival NVIDIA's NVLink.
- Thermal Design Power (TDP): Rated at 550 W.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (24)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- nvidia.com
- deepinfra.com
- siliconflow.com
- huggingface.co
- openrouter.ai
- deepinfra.com
- huaweicentral.com
- heim.xyz
- eeworld.com.cn
- tomshardware.com
- huaweicentral.com
- seekingalpha.com
- scmp.com
- emergentmind.com
- benchgecko.ai
- substack.com
- georgetown.edu
- youtube.com
- medium.com
- reddit.com
- kili-technology.com
- laweconcenter.org
- ai-frontiers.org
- merics.org
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology โ
