China launches GPU-free LineShine supercomputer with 2.45M CPU cores

๐กSee how China is scaling exascale computing without GPUs, potentially shifting the future of AI infrastructure.
โก 30-Second TL;DR
What Changed
System delivers 1.54 exaFLOPS performance using only CPU architecture.
Why It Matters
This development highlights a significant push to achieve computational sovereignty by decoupling high-performance AI and scientific workloads from reliance on Western GPU supply chains.
What To Do Next
Evaluate whether your current model training or inference pipelines can be optimized for CPU-based clusters to reduce dependency on scarce GPU resources.
Key Points
- โขSystem delivers 1.54 exaFLOPS performance using only CPU architecture.
- โขPowered by 2.45 million domestic LX2 processors based on the Armv9 instruction set.
- โขSignals a strategic shift toward CPU-centric high-performance computing to bypass GPU supply constraints.
๐ง Deep Insight
Web-grounded analysis with 12 cited sources.
๐ Enhanced Key Takeaways
- โขThe LX2 processor, potentially designed by Huawei, is a custom Armv9-based CPU featuring 304 cores per processor, an unusual memory subsystem with 32 GB of on-package HBM (4 TB/s bandwidth), and up to 256 GB of off-package DDR5 memory.
- โขLineShine achieves 1.54 exaFLOPS in BF16 training performance and can peak at 2.16 exaFLOPS when training a 6.3-billion-parameter Earth observation generative compression model.
- โขThe supercomputer is interconnected by China's proprietary LingQi (LQLink) high-speed network, which uses a dual-plane multi-rail fat-tree topology and provides 1.6 Tb/s bandwidth per node.
- โขThe system runs on Anolis OS 8.9, an Alibaba-developed RHEL-compatible Linux distribution, and includes a ROCm-compatible environment, GCC, rocBLAS, and PyTorch 2.7.1.
- โขLineShine is explicitly positioned as a direct response to US export controls on advanced GPUs, demonstrating China's commitment to achieving 'full-stack independence' in high-performance computing through entirely domestic hardware and software.
๐ Competitor Analysisโธ Show
| Feature/Benchmark | LineShine (China) | El Capitan (USA) | Frontier (USA) | Aurora (USA) |
|---|---|---|---|---|
| Architecture | CPU-only (Armv9-based LX2) | CPU+GPU (AMD EPYC + AMD Instinct MI300A) | CPU+GPU (AMD EPYC + AMD Instinct MI250X) | CPU+GPU (Intel Xeon Max + Intel Data Center Max GPUs) |
| Peak Performance | 2.16 ExaFLOPS (BF16 training) / Claimed >2 ExaFLOPS (sustained FP64) | 2.79 ExaFLOPS (peak) | 1.353 ExaFLOPS (Rmax) | 1.012 ExaFLOPS (Rmax) |
| Sustained Performance | 1.54 ExaFLOPS (BF16 training) | 1.742 ExaFLOPS (Rmax Linpack) | 1.353 ExaFLOPS (Rmax Linpack) | 1.012 ExaFLOPS (Rmax Linpack) |
| Processor Count | 40,960 LX2 processors (2.45M cores) | >11,000 nodes (over 11M cores) | N/A | N/A |
| Interconnect | LingQi (LQLink), 1.6 Tb/s per node | Slingshot 11 | HPE Cray Slingshot | N/A |
| Memory Subsystem | On-package HBM + off-package DDR5 | HBM3 | HBM2e | HBM2e |
| Domestic Origin | Full-stack domestic (chips, network, storage, OS) | US-built | US-built | US-built |
| TOP500 Status | No independent Linpack data published; not expected to appear on TOP500 | #1 as of November 2025 (1.742 ExaFLOPS Rmax) | #2 as of November 2024 (1.353 ExaFLOPS Rmax) | #3 as of November 2024 (1.012 ExaFLOPS Rmax) |
๐ ๏ธ Technical Deep Dive
- Processor (LX2): Armv9-based, with each processor integrating two compute chiplets for a total of 304 CPU cores. These cores are organized into eight CPU clusters, each containing 38 cores.
- Memory Architecture: Each LX2 processor features a hybrid memory subsystem, combining 32 GB of on-package High Bandwidth Memory (HBM) providing approximately 4 TB/s of bandwidth, and up to 256 GB of off-package DDR5 memory. The processor has 16 NUMA domains, with four HBM and four DDR domains per chiplet. A dedicated SDMA engine manages data movement between DDR and HBM.
- Node Configuration: The LineShine system comprises 20,480 compute nodes, with each node housing two LX2 processors, totaling 608 cores, 64 GB HBM, and 512 GB DDR per node.
- Interconnect: The LingQi (LQLink) high-speed network utilizes a dual-plane multi-rail fat-tree topology, delivering 1.6 Tb/s bandwidth per node.
- Performance Capabilities per LX2: A single LX2 processor delivers 60.3 TFLOPS in FP64 performance, 240 TFLOPS in BF16/FP16 throughput, and 960 TOPS in INT8 performance, optimized for dense AI and matrix workloads via SME and SVE vector units.
- Storage System: The supercomputer includes 650 PB of storage capacity with an aggregate bandwidth of 10 TB/s, distributed across 428 storage nodes in 67 liquid-cooled cabinets.
- Operating System & Software Stack: LineShine runs on Anolis OS 8.9, an Alibaba-developed RHEL-compatible Linux distribution. Its software environment is ROCm-compatible and includes GCC 8.5.0, rocBLAS, and PyTorch 2.7.1, alongside a software-defined asynchronous MPI runtime for optimized performance.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (12)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: TechNode โ