HK's DeepSeek Model Runs on China Chips Abroad

๐กSovereign LLM fully on China chips launches abroadโkey for infra sovereignty tests.
โก 30-Second TL;DR
What Changed
HKGAI-V3 based on DeepSeek V4 with full-parameter fine-tuning
Why It Matters
Advances China's AI independence by enabling advanced models on domestic chips, potentially lowering costs and geopolitical risks for users. Positions Hong Kong as a hub for exporting China-centric AI solutions globally.
What To Do Next
Monitor HKGAI site for HKGAI-V3 H1 release and benchmark on Chinese chips vs DeepSeek.
Key Points
- โขHKGAI-V3 based on DeepSeek V4 with full-parameter fine-tuning
- โขOptimized to run entirely on Chinese-made chips
- โขGovernment-backed lab targeting sovereign AI exports
- โขPlanned unveiling in first half of this year
๐ง Deep Insight
Web-grounded analysis with 3 cited sources.
๐ Enhanced Key Takeaways
- โขDeepSeek V4, the underlying architecture for HKGAI-V3, was released on April 24, 2026, featuring a 1.6 trillion-parameter Mixture-of-Experts (MoE) design and a 1-million-token context window.
- โขThe initiative is part of a broader 'de-CUDA-fy' strategy in China, where DeepSeek provided exclusive early optimization access to Huawei and other domestic chipmakers, intentionally excluding Nvidia and AMD from pre-release hardware tuning.
- โขHKGAI's sovereign AI approach focuses on 'governance-embedded' alignment, specifically addressing Hong Kong's unique multilingual (Cantonese, Mandarin, English) and socio-legal requirements under the 'one country, two systems' framework.
๐ Competitor Analysisโธ Show
| Feature | DeepSeek V4-Pro | GPT-5.5 | Gemini 3.1-Pro |
|---|---|---|---|
| Architecture | 1.6T MoE | Proprietary | Proprietary |
| Context Window | 1M tokens | 1M+ tokens | 1M+ tokens |
| Pricing (Input/Output per 1M) | $1.74 / $3.48 | $5.00 / $30.00 | N/A (Enterprise) |
| Hardware Optimization | Huawei Ascend 950 | Nvidia H100/B200 | Google TPU |
| Benchmark (Reasoning) | 52 (Index) | 60 (Index) | N/A |
๐ ๏ธ Technical Deep Dive
- Model Architecture: DeepSeek V4-Pro utilizes a 1.6 trillion-parameter Mixture-of-Experts (MoE) architecture with 49 billion active parameters per token.
- Training/Inference: Employs FP4+FP8 mixed-precision training. Inference is specifically optimized for Huawei Ascend 950 supernode clusters.
- Performance: V4-Pro achieves a single-card decode throughput of 4,700 tokens per second (TPS) on Ascend 950 hardware for 8K input scenarios.
- Alignment: HKGAI-V1 (predecessor) utilized full-parameter fine-tuning to instantiate regional values, using bespoke benchmarks like HKMMLU (local knowledge), SafeLawBench, and Adversarial HK Value Bench.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (3)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology โ