🐯Freshcollected in 17m

Domestic GPUs Enter the Big-Tech Finals

Domestic GPUs Enter the Big-Tech Finals
PostLinkedIn
🐯Read original on 虎嗅

💡China’s GPU race is moving from policy-driven orders to capacity battles and internet-scale validation.

⚡ 30-Second TL;DR

What Changed

Cambricon reported first-half revenue of RMB 5.996 billion and net profit of RMB 2.311 billion, both roughly doubling year over year.

Why It Matters

For AI infrastructure founders, access to advanced-node capacity may be as decisive as chip architecture. Vendors without reliable 7nm-class supply and proven software integration risk being confined to smaller inference deployments, while internet customer validation becomes the main route to recurring revenue.

What To Do Next

Benchmark your AI accelerator stack against representative training and inference workloads, then document model compatibility, token cost, cluster scaling, and supply commitments before approaching internet customers.

Who should care:Founders & Product Leaders

Key Points

  • Cambricon reported first-half revenue of RMB 5.996 billion and net profit of RMB 2.311 billion, both roughly doubling year over year.
  • Moore Threads and Enflame are pursuing scale, while companies including Moore Threads and Cambricon are building inventories and prepaying suppliers to secure capacity.
  • SMIC is described as the only domestic foundry currently mass-producing 7nm-class AI chips at scale, creating a severe supply bottleneck.
  • Huawei reportedly receives about 43% of relevant capacity, while Cambricon receives approximately 9%-11%; other GPU vendors compete for the remainder.
  • Internet customers evaluate hardware, model compatibility, and cluster performance, with the path from sample testing to mass procurement potentially taking two years or longer.

🧠 Deep Insight

AI-generated analysis for this event.

🔑 Enhanced Key Takeaways

  • The US export controls on high-end AI chips (such as NVIDIA's A100/H100 and their China-specific variants) have acted as the primary catalyst for the rapid shift toward domestic GPU adoption in Chinese data centers.
  • Major Chinese cloud providers, including Alibaba Cloud, Baidu, and ByteDance, have established dedicated internal teams to optimize software stacks (like CUDA-to-CANN or custom operator libraries) to mitigate the performance gap of domestic hardware.
  • The 'wafer bottleneck' is exacerbated by the high demand for HBM (High Bandwidth Memory) and advanced packaging (CoWoS-like) capacity, which remains a significant hurdle for domestic GPU manufacturers compared to global leaders.
  • Financial reports indicate that while revenue is surging, R&D expenditure remains extremely high, often exceeding 50% of total revenue, as companies race to achieve software ecosystem parity with NVIDIA's CUDA platform.
  • Recent industry data suggests a consolidation trend where smaller, less-funded GPU startups are struggling to secure foundry slots at SMIC, leading to potential M&A activity within the domestic semiconductor sector.
📊 Competitor Analysis▸ Show
FeatureCambricon (MLU Series)Moore Threads (MTT Series)NVIDIA (H20/L20)
Primary FocusTraining & InferenceGeneral Purpose GPUInference/Training
Software StackCambricon NeuwareMUSACUDA
Process Node7nm (SMIC)7nm (SMIC)7nm (TSMC)
Market PositionEstablished/PublicEmerging/PrivateIncumbent (Restricted)

🛠️ Technical Deep Dive

  • Cambricon MLU series utilizes a proprietary architecture optimized for tensor operations, often featuring high-speed interconnects designed to scale across multi-node clusters.
  • Moore Threads MUSA architecture supports a wide range of precision formats including FP32, FP16, and INT8, with specific hardware acceleration for video encoding/decoding and AI inference tasks.
  • Domestic chips are increasingly relying on 2.5D packaging technologies to integrate HBM or high-speed GDDR6 memory, which is critical for overcoming the memory wall in large language model (LLM) training.
  • Software compatibility layers are being developed to allow existing PyTorch and TensorFlow models to run on domestic hardware with minimal code changes, though performance overhead remains a key technical challenge.

🔮 Future ImplicationsAI analysis grounded in cited sources

Domestic GPU market share in Chinese data centers will exceed 40% by 2027.
The combination of forced supply chain localization and aggressive software optimization by major internet firms is rapidly closing the usability gap.
SMIC will face a capacity crisis for 7nm-class AI chips by mid-2027.
The current demand from multiple domestic GPU vendors and Huawei's own internal requirements far outstrips the current yield-limited capacity of SMIC's advanced nodes.

Timeline

2016-03
Cambricon Technologies is founded as a spin-off from the Chinese Academy of Sciences.
2020-11
Cambricon completes its IPO on the Shanghai Stock Exchange STAR Market.
2022-10
US government implements sweeping export controls on advanced AI chips to China, accelerating domestic substitution efforts.
2024-05
Cambricon reports significant revenue growth driven by increased demand for AI training clusters from domestic internet giants.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅

Domestic GPUs Enter the Big-Tech Finals | 虎嗅 | SetupAI | SetupAI