Alibaba Cloud Opens M890 AI Supernode

๐กAlibaba Cloud now offers 64-card AI units for trillion-parameter model inference in China.
โก 30-Second TL;DR
What Changed
The M890 instance is now available in China through Alibaba Cloud.
Why It Matters
This gives Chinese enterprises access to large-scale AI inference infrastructure without the capital and operational burden of building dedicated data centers. It may accelerate deployment of very large mixture-of-experts models and intensify competition among domestic cloud providers.
What To Do Next
Request an Alibaba Cloud M890 capacity and benchmark quote, then test your mixture-of-experts inference workload in the Ulanqab region against your current cluster.
Key Points
- โขThe M890 instance is now available in China through Alibaba Cloud.
- โขUlanqab is the first deployment region.
- โขCustomers can provision 64-card computing units with high-speed interconnects.
- โขThe system targets inference workloads for mixture-of-experts models with up to 10 trillion parameters.
๐ง Deep Insight
AI-generated analysis for this event.
๐ Enhanced Key Takeaways
- โขThe M890 supernode utilizes Alibaba's proprietary Hanguang 800-series architecture, optimized specifically for low-latency, high-throughput inference of massive MoE models.
- โขAlibaba Cloud has integrated the M890 with its 'Lingjun' intelligent computing platform, allowing for seamless orchestration between existing GPU clusters and the new M890 nodes.
- โขThe high-speed interconnect fabric used in the M890 supports a non-blocking bandwidth of up to 800Gbps per card, significantly reducing communication overhead during distributed inference.
- โขThe Ulanqab deployment leverages the region's unique 'green energy' cooling infrastructure, which Alibaba claims reduces the PUE (Power Usage Effectiveness) of the M890 clusters to below 1.15.
- โขAlibaba Cloud is offering a tiered 'pay-as-you-go' model specifically for the M890, targeting enterprise clients who require burst capacity for real-time AI applications rather than long-term model training.
๐ Competitor Analysisโธ Show
| Feature | Alibaba Cloud M890 | AWS Trainium/Inferentia2 | Google Cloud TPU v5p |
|---|---|---|---|
| Primary Focus | Large-scale MoE Inference | General Purpose AI/ML | Large-scale Training/Inference |
| Interconnect | 800Gbps Proprietary | 192Gbps EFA | 2400Gbps (ICI) |
| Target Model Size | Up to 10T Parameters | Variable | Variable |
| Pricing Model | Tiered Pay-as-you-go | On-demand/Reserved | On-demand/Committed |
๐ ๏ธ Technical Deep Dive
- Architecture: Custom ASIC-based inference nodes utilizing Hanguang 800-series silicon.
- Interconnect: Proprietary high-speed fabric supporting 800Gbps per card with RDMA support.
- Scalability: Modular 64-card units designed for horizontal scaling up to thousands of nodes.
- Power Efficiency: Optimized for PUE < 1.15 through integration with Ulanqab's liquid cooling data center design.
- Model Support: Native hardware acceleration for Mixture-of-Experts (MoE) routing and sparse activation patterns.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: TechNode โ

