Moore Threads Joins National Embodied AI Pilot Base
💡Moore Threads is positioning its GPUs as a key infrastructure for China's growing embodied AI and robotics sector.
⚡ 30-Second TL;DR
What Changed
Moore Threads joins the National Embodied AI Pilot Base as a partner.
Why It Matters
This partnership strengthens the domestic hardware ecosystem for robotics and embodied AI, potentially accelerating the training and simulation capabilities for Chinese humanoid robot developers.
What To Do Next
If you are developing robotics software, monitor the Moore Threads developer portal for upcoming SDK updates related to their embodied AI simulation tools.
Key Points
- •Moore Threads joins the National Embodied AI Pilot Base as a partner.
- •Establishment of a joint laboratory for embodied AI computing and simulation.
- •Moore Threads will contribute to the industrial committee of the base.
🧠 Deep Insight
Web-grounded analysis with 12 cited sources.
🔑 Enhanced Key Takeaways
- •The partnership with the National Embodied AI Pilot Base aligns with China's national strategy to cultivate a domestic AI ecosystem and reduce reliance on foreign AI computing and robotics software, particularly in light of US export controls.
- •The newly established joint laboratory will combine Moore Threads' full-function GPU platform and KUAE AI computing cluster with Lightwheel.ai's simulation framework to generate high-fidelity synthetic data crucial for training embodied AI and robotics systems.
- •Embodied AI is a national priority for China, explicitly included in the 15th Five-Year Plan (2026-2030), with the goal of integrating AI into physical systems like robots, drones, and autonomous vehicles to enhance productivity and potentially achieve Artificial General Intelligence (AGI).
- •The National Embodied AI Pilot Base, launched in Hangzhou, operates with over 130 robots across more than 30 vocational scenarios, aiming to foster collaboration and coordinated development in foundational chips, operating systems, and real-world AI applications.
- •Moore Threads' MUSA architecture is designed to simultaneously support diverse workloads including AI computing, graphics rendering, scientific computing, physical simulation, and ultra-high-definition video encoding/decoding, making its GPUs versatile for embodied AI applications.
📊 Competitor Analysis▸ Show
| Company | Key AI Products/Focus | Performance (relative to NVIDIA) | Software Ecosystem | Other Notes |
|---|---|---|---|---|
| Moore Threads | MTT S4000, Huashan (next-gen AI GPU) | MTT S4000: 25 TFLOPS FP32, 200 TFLOPS FP16/BF16. Huashan: Claims compute density +50%, energy efficiency +10x, memory bandwidth rivaling/exceeding Blackwell B200. | MUSA, MUSIFY (CUDA compatibility) | Broadest use case coverage (gaming, AI, simulation). |
| Huawei HiSilicon | Ascend 910B/C | Ascend 910B: ~80% of A100 performance, 4 PetaFLOPS AI performance. | CANN (Huawei's AI computing architecture) | Largest installed base among Chinese alternatives (200K+ units shipped). |
| Cambricon | MLU590, Cambricon-1M (edge), Cambricon 290 (data center) | MLU590: 60-70% of A100 performance. Cambricon-1M: 4 TOPS at 2W. | BANG (Cambricon's AI software platform) | Specialized AI processor innovator for edge and cloud. Publicly listed (SH: 688256). |
| Biren Technology | BR100 | Targets H100-level performance. | BIRENSUPA | Faced supply chain challenges. |
| Enflame (Suiyuan) | GCU series | Cloud-focused AI accelerator. | Strong inference optimization, backed by Tencent. |
🛠️ Technical Deep Dive
- MUSA Architecture: Moore Threads Unified System Architecture, with the third-generation powering the MTT S4000 and the next-gen 'Huagang' architecture for Lushan (gaming) and Huashan (AI) GPUs.
- MTT S4000 Accelerator: Features 8192 MUSA Cores and 128 Tensor Cores. Equipped with 48GB of VRAM and 768GB/s memory bandwidth. Supports FP64, FP32 (25 TFLOPS), TF32 (50 TFLOPS), FP16/BF16 (200 TFLOPS), and INT8 (200 TOPS) precision computing. Utilizes PCIe 5.0 x16 interface and MTLink 1.0 for multi-GPU interconnectivity, enabling clusters with thousands of cards. Includes MUSA Security Engine 2.0 and supports hardware virtualization, GPU elastic partitioning, and SR-IOV isolation.
- Huashan AI GPU (Next-Gen): Designed with a chiplet-based architecture, incorporating two compute dies and eight stacks of high-bandwidth memory. Leverages MTLink 4.0 interconnect technology to scale AI training clusters to over 100,000 GPUs. Claims a 50% increase in compute density and up to a 10x improvement in energy efficiency compared to earlier designs. Memory bandwidth is touted to rival or exceed NVIDIA's Blackwell B200. Supports a wide range of compute formats from FP4 to FP64, including exclusive low-precision mixed formats like MTFP4, MTFP6, and MTFP8.
- KUAE Intelligent Computing Center: A 1000-card AI training cluster built around the MTT S4000 GPUs, integrating RDMA networking, distributed storage, and comprehensive cluster management software. The KUAE Platform manages multi-datacenter resource allocation, and KUAE ModelStudio hosts training frameworks and model repositories. Demonstrated near-linear 91% scaling efficiency.
- Software Ecosystem: Moore Threads provides the MUSA software stack and a tool called MUSIFY, which reportedly translates CUDA code to the MUSA GPU architecture with zero migration cost or performance penalty.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (12)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗
