๐ŸคStalecollected in 34h

Together GPU Clusters Add Autoscaling & Self-Healing

Together GPU Clusters Add Autoscaling & Self-Healing
PostLinkedIn
๐ŸคRead original on Together AI Blog
#autoscaling#observability#self-healing#rbactogether-gpu-clusterstogether-ai

๐Ÿ’กProduction GPU clusters with autoscaling & self-healing cut AI infra costs 30-50%.

โšก 30-Second TL;DR

What Changed

Built-in autoscaling for dynamic resource scaling

Why It Matters

This reduces ops overhead for AI teams, enabling focus on development over maintenance. It improves cost-efficiency and reliability for large-scale ML training and inference.

What To Do Next

Sign up for Together GPU Clusters and test autoscaling on a sample ML workload.

Who should care:Enterprise & Security Teams

Key Points

  • โ€ขBuilt-in autoscaling for dynamic resource scaling
  • โ€ขRBAC for secure role-based access control
  • โ€ขFull-stack observability for comprehensive monitoring
  • โ€ขSelf-healing node repair for automatic failure recovery

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 8 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขTogether AI is reportedly pursuing a $1 billion funding round amid these infrastructure upgrades.[1]
  • โ€ขAutoscaling is powered by Kubernetes Cluster Autoscaler, which monitors GPU-constrained workloads and automatically provisions or decommissions nodes based on real-time demand.[1][2]
  • โ€ขSelf-healing involves a three-click process where the control plane cords, drains, and recreates failed nodes, with automatic acceptance tests ensuring clusters pass before marking as ready.[2]
  • โ€ขGPU Clusters support NVIDIA H100, H200, B200, and GB200 GPUs with InfiniBand networking, scaling from 8 to over 4,000 GPUs for distributed training and inference.[3]

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขAutoscaling uses Kubernetes Cluster Autoscaler for dynamic node provisioning/deprovisioning based on GPU demand, targeting variable inference and bursty training jobs.[1][2]
  • โ€ขObservability features a dedicated Grafana instance with dashboards for DCGM GPU metrics, InfiniBand/NIC networking telemetry, storage I/O, and Kubernetes health.[2]
  • โ€ขSelf-healing node repair: Cordon and drain failed nodes, recreate on new/existing hosts, run acceptance tests (detailed in documentation) before readiness.[2]
  • โ€ขOrchestration options include managed Kubernetes (Kubeadm-based, HA control plane, node autoscaling, Cert Manager) and Slurm on Kubernetes for precise gang scheduling and SSH access.[3]
  • โ€ขPerformance enhanced by Together Kernel Collection (including ThunderKittens), achieving 90% faster training on B200 (15,264 tokens/s/GPU) vs H100 (8,080 tokens/s/GPU) for 70B Llama model.[3]

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Together GPU Clusters will capture more enterprise AI workloads from static provisioning providers.
Dynamic autoscaling and self-healing address key pain points of over/underprovisioning and node failures, enabling efficient scaling for variable training and inference.[1][2]
$1B funding pursuit will accelerate Together AI's GPU infrastructure expansion.
The upgrades coincide with reported funding talks, providing production-ready features to attract enterprise customers and investors in the competitive GPU-as-a-service market.[1][7]

โณ Timeline

2025-03
Introduced Instant GPU Clusters at NVIDIA GTC 2025.
2025-09
Launched self-service GPU infrastructure with Together Instant Clusters.
2026-03
Announced autoscaling, RBAC, observability, and self-healing for GPU Clusters.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Together AI Blog โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.