Hong Kong launches DeepSeek-based model for domestic chips

๐กFirst DeepSeek-based model optimized for domestic chips, offering a path to bypass high-end GPU supply constraints.
โก 30-Second TL;DR
What Changed
HKGAI-V3 is built upon the DeepSeek V4 architecture.
Why It Matters
This development signals a strategic shift toward hardware-software co-optimization in the Chinese AI ecosystem, reducing reliance on foreign GPU architectures.
What To Do Next
Evaluate the HKGAI-V3 model if you are developing AI applications for deployment on domestic Chinese hardware infrastructure.
Key Points
- โขHKGAI-V3 is built upon the DeepSeek V4 architecture.
- โขThe model is specifically optimized to run on domestic Chinese AI chips.
- โขAchieved a tenfold improvement in token efficiency and enhanced agentic capabilities.
๐ง Deep Insight
Web-grounded analysis with 15 cited sources.
๐ Enhanced Key Takeaways
- โขHKGAI-V3 is designed with embedded local cultural context and supports Cantonese, alongside Chinese and English, to improve comprehension and outputs for city-specific use cases in Hong Kong.
- โขThe model's Agent Workshop platform demonstrated a near hundredfold increase in uninterrupted agent runtime, operating stably for up to 28 hours in a single session for tasks such as generating research reports.
- โขHKGAI-V3 is optimized for both mainstream and domestic Chinese hardware, specifically mentioning compatibility with Huawei Technologies' Ascend 910C chips.
- โขThis initiative is a key component of Hong Kong's broader 'sovereign AI' strategy, which aims for the jurisdiction to develop and operate AI systems on its own infrastructure to ensure security and data sovereignty.
- โขThe Hong Kong Generative AI Research and Development Centre (HKGAI Centre), established in October 2023 under the InnoHK programme, previously developed HKChat, the city's first Cantonese-enabled chatbot for local services and regulations.
๐ ๏ธ Technical Deep Dive
- DeepSeek V4 is a Mixture-of-Experts (MoE) model, with the Pro version featuring 1.6 trillion total parameters (49 billion active) and the Flash version having 284 billion total parameters (13 billion active).
- It supports an ultra-long context window of up to 1 million tokens, equivalent to 15-20 full novels.
- Key architectural innovations include a Hybrid Attention Architecture (combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)) for enhanced long-context efficiency.
- The model incorporates Manifold-Constrained Hyper-Connections (mHC) to stabilize signal propagation across deep layers and utilizes the Muon Optimizer for faster convergence and improved training stability.
- DeepSeek V4 employs mixed precision training, with MoE expert parameters using FP4 precision and most other parameters using FP8 to maximize memory efficiency.
- The V4 model has been validated across both Nvidia Graphics Processing Units (GPUs) and Huawei Ascend Neural Processing Units (NPUs), indicating a fine-grained expert parallelism scheme that supports both platforms.
- HKGAI-V3 involves "full-parameter fine tuning for localisation" of the DeepSeek V4 architecture.
- DeepSeek V4 also features native multimodal support, processing text, images, video, and audio from scratch.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology โ