Four M5 Ultra Mac Studios Deliver 4.8TB/s Bandwidth
💡A four-Mac cluster claims data-center-class bandwidth for local execution of giant LLMs.
⚡ 30-Second TL;DR
What Changed
EXO Labs demonstrated a four-Mac Studio cluster built around the M5 Ultra.
Why It Matters
A compact multi-Mac cluster could give researchers and developers an alternative to cloud GPU infrastructure for certain large-model workloads. Real-world value will depend on interconnect overhead, model partitioning efficiency, software support, power consumption, and total cost.
What To Do Next
Benchmark your target Kimi K3 workload on a four-node M5 Ultra cluster, comparing tokens per second, power draw, and total cost against your current cloud GPU setup.
Key Points
- •EXO Labs demonstrated a four-Mac Studio cluster built around the M5 Ultra.
- •The cluster reportedly reaches approximately 4.8TB/s of aggregate memory bandwidth.
- •The system is positioned for local execution of very large models such as Kimi K3.
🧠 Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
🔑 Enhanced Key Takeaways
- •The cluster architecture relies on a proprietary RDMA implementation over Thunderbolt 5, developed through a year-long partnership between EXO Labs and Apple.
- •Each individual M5 Ultra chip provides 1.2TB/s of memory bandwidth, representing a 50% increase over the preceding M3 Ultra generation.
- •The M5 Ultra utilizes a quad-die design, integrating two M5 Max chips via UltraFusion to support up to 512GB of unified memory.
- •The new Mac Studio hardware introduces support for Wi-Fi 7, Bluetooth 6, and 120Gb/s Thunderbolt 5 connectivity to facilitate high-speed clustering.
- •Performance benchmarks indicate the M5 Ultra delivers 4.3x faster AI-specific processing compared to the M3 Ultra, alongside 1.8x faster graphics performance.
📊 Competitor Analysis▸ Show
| Feature | M5 Ultra Cluster (4-Node) | NVIDIA DGX Station A100 | Enterprise Workstation (Dual RTX 6000 Ada) |
|---|---|---|---|
| Aggregate Bandwidth | 4.8 TB/s | ~2.0 TB/s | ~1.9 TB/s |
| Unified Memory | 2TB (512GB x 4) | 320GB | 96GB |
| Connectivity | Thunderbolt 5 RDMA | InfiniBand/Ethernet | PCIe 5.0 |
| Pricing (Approx) | ~3.8M JPY | ~$150,000+ | ~$15,000 |
🛠️ Technical Deep Dive
- Architecture: Quad-die design utilizing UltraFusion interconnects to bridge two M5 Max chips.
- Memory: Unified memory architecture supporting up to 512GB per node.
- Networking: Implementation of low-latency RDMA (Remote Direct Memory Access) specifically optimized for the 120Gb/s Thunderbolt 5 interface.
- Compute: 36-core CPU and 80-core GPU configuration per M5 Ultra chip.
- Throughput: 1.2TB/s memory bandwidth per chip, enabling the 4.8TB/s aggregate figure in a 4-node configuration.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.

