Hong Kong launches DeepSeek-based model for domestic chips

💡First DeepSeek-based model optimized for domestic chips, offering a path to bypass high-end GPU supply constraints.
⚡ 30-Second TL;DR
What Changed
HKGAI-V3 is built upon the DeepSeek V4 architecture.
Why It Matters
This development signals a strategic shift toward hardware-software co-optimization in the Chinese AI ecosystem, reducing reliance on foreign GPU architectures.
What To Do Next
Evaluate the HKGAI-V3 model if you are developing AI applications for deployment on domestic Chinese hardware infrastructure.
Key Points
- •HKGAI-V3 is built upon the DeepSeek V4 architecture.
- •The model is specifically optimized to run on domestic Chinese AI chips.
- •Achieved a tenfold improvement in token efficiency and enhanced agentic capabilities.
🧠 Deep Insight
Background and context from public sources — not the original article. 15 sources cited.
🔑 Enhanced Key Takeaways
- •HKGAI-V3 is designed with embedded local cultural context and supports Cantonese, alongside Chinese and English, to improve comprehension and outputs for city-specific use cases in Hong Kong.
- •The model's Agent Workshop platform demonstrated a near hundredfold increase in uninterrupted agent runtime, operating stably for up to 28 hours in a single session for tasks such as generating research reports.
- •HKGAI-V3 is optimized for both mainstream and domestic Chinese hardware, specifically mentioning compatibility with Huawei Technologies' Ascend 910C chips.
- •This initiative is a key component of Hong Kong's broader 'sovereign AI' strategy, which aims for the jurisdiction to develop and operate AI systems on its own infrastructure to ensure security and data sovereignty.
- •The Hong Kong Generative AI Research and Development Centre (HKGAI Centre), established in October 2023 under the InnoHK programme, previously developed HKChat, the city's first Cantonese-enabled chatbot for local services and regulations.
🛠️ Technical Deep Dive
- DeepSeek V4 is a Mixture-of-Experts (MoE) model, with the Pro version featuring 1.6 trillion total parameters (49 billion active) and the Flash version having 284 billion total parameters (13 billion active).
- It supports an ultra-long context window of up to 1 million tokens, equivalent to 15-20 full novels.
- Key architectural innovations include a Hybrid Attention Architecture (combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)) for enhanced long-context efficiency.
- The model incorporates Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across deep layers and utilizes the Muon Optimizer for faster convergence and improved training stability.
- DeepSeek V4 employs mixed precision training, with MoE expert parameters using FP4 precision and most other parameters using FP8 to maximize memory efficiency.
- The V4 model has been validated across both Nvidia Graphics Processing Units (GPUs) and Huawei Ascend Neural Processing Units (NPUs), indicating a fine-grained expert parallelism scheme that supports both platforms.
- HKGAI-V3 involves "full-parameter fine tuning for localisation" of the DeepSeek V4 architecture.
- DeepSeek V4 also features native multimodal support, processing text, images, video, and audio from scratch.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.
