IBM Lands $240M Together AI Compute Deal

💡A major $240M B300 deployment could reshape open-source model inference capacity and cloud options.
⚡ 30-Second TL;DR
What Changed
The multiyear agreement is valued at $240 million.
Why It Matters
The deal expands cloud-based capacity for production open-source model inference and signals growing enterprise demand for high-end AI infrastructure. It may give Together AI more room to serve large workloads while keeping open models competitive with closed systems.
What To Do Next
Add the announced IBM Cloud HGX B300 capacity to your 2027 inference roadmap and compare its expected cost and latency with current GPU providers.
Key Points
- •The multiyear agreement is valued at $240 million.
- •IBM Cloud will host an NVIDIA HGX B300 cluster dedicated to large-scale inference.
- •The deployment combines NVIDIA HGX B300 systems with Spectrum-X Ethernet technology.
- •Together AI recently raised $800 million in Series C funding at an $8.3 billion valuation.
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •The NVIDIA HGX B300 cluster utilizes the Blackwell architecture, specifically optimized for high-throughput inference workloads rather than just training.
- •IBM's integration of Spectrum-X Ethernet is a strategic move to provide AI-optimized networking that avoids the complexity and cost of InfiniBand for large-scale inference clusters.
- •This deal marks a significant expansion of IBM's 'AI-as-a-Service' strategy, positioning IBM Cloud as a specialized provider for open-source model providers rather than just enterprise software clients.
- •Together AI plans to leverage this infrastructure to reduce latency for its API-based inference services, directly challenging proprietary model providers like OpenAI and Anthropic.
- •The partnership includes a collaborative engineering component where IBM and Together AI will co-optimize the software stack to improve GPU utilization rates on the B300 hardware.
📊 Competitor Analysis▸ Show
| Feature | IBM Cloud + Together AI | AWS (Bedrock/Trainium) | Google Cloud (TPU/Vertex AI) |
|---|---|---|---|
| Primary Hardware | NVIDIA HGX B300 | Trainium2 / H200 | TPU v5p / H100 |
| Networking | Spectrum-X Ethernet | EFA (Elastic Fabric Adapter) | Jupiter / Custom Interconnect |
| Model Focus | Open-Source Inference | Proprietary & Open | Proprietary & Open |
| Pricing Model | Dedicated Cluster/Reserved | On-demand/Provisioned | On-demand/Reserved |
🛠️ Technical Deep Dive
- The NVIDIA HGX B300 platform integrates eight Blackwell GPUs per node, connected via NVLink for high-bandwidth memory access.
- Spectrum-X Ethernet technology provides 400Gb/s or 800Gb/s throughput with adaptive routing and congestion control specifically tuned for AI traffic patterns.
- The deployment utilizes IBM's VPC (Virtual Private Cloud) architecture to ensure multi-tenant isolation while maintaining bare-metal performance for the inference cluster.
- Together AI's inference stack is optimized for the Blackwell architecture's Transformer Engine, which supports FP4 and FP6 precision to accelerate token generation speeds.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: IT之家 ↗


