🤝較早收集於 34h

Together GPU Clusters 新增自動擴展與自癒

Together GPU Clusters 新增自動擴展與自癒
PostLinkedIn
🤝閱讀原文: Together AI Blog
#autoscaling#observability#self-healing#rbactogether-gpu-clusterstogether-ai

💡具自動擴展與自癒的生產 GPU 叢集,可降低 AI 基礎設施成本 30-50%。

⚡ 30-Second TL;DR

有什麼變化

內建自動擴展實現動態資源調整

為什麼重要

這降低 AI 團隊的運維負擔,讓他們專注開發而非維護。提升大規模 ML 訓練與推論的成本效益與可靠性。

下一步行動

註冊 Together GPU Clusters,並在範例 ML 工作負載上測試自動擴展。

誰應關注:Enterprise & Security Teams

關鍵要點

  • 內建自動擴展實現動態資源調整
  • RBAC 提供安全角色存取控制
  • 全棧觀測性實現全面監控
  • 自癒節點修復自動故障恢復

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 8 個來源。

🔑 增強重點摘要

  • Together AI is reportedly pursuing a $1 billion funding round amid these infrastructure upgrades.[1]
  • Autoscaling is powered by Kubernetes Cluster Autoscaler, which monitors GPU-constrained workloads and automatically provisions or decommissions nodes based on real-time demand.[1][2]
  • Self-healing involves a three-click process where the control plane cords, drains, and recreates failed nodes, with automatic acceptance tests ensuring clusters pass before marking as ready.[2]
  • GPU Clusters support NVIDIA H100, H200, B200, and GB200 GPUs with InfiniBand networking, scaling from 8 to over 4,000 GPUs for distributed training and inference.[3]

🛠️ 技術深入

  • Autoscaling uses Kubernetes Cluster Autoscaler for dynamic node provisioning/deprovisioning based on GPU demand, targeting variable inference and bursty training jobs.[1][2]
  • Observability features a dedicated Grafana instance with dashboards for DCGM GPU metrics, InfiniBand/NIC networking telemetry, storage I/O, and Kubernetes health.[2]
  • Self-healing node repair: Cordon and drain failed nodes, recreate on new/existing hosts, run acceptance tests (detailed in documentation) before readiness.[2]
  • Orchestration options include managed Kubernetes (Kubeadm-based, HA control plane, node autoscaling, Cert Manager) and Slurm on Kubernetes for precise gang scheduling and SSH access.[3]
  • Performance enhanced by Together Kernel Collection (including ThunderKittens), achieving 90% faster training on B200 (15,264 tokens/s/GPU) vs H100 (8,080 tokens/s/GPU) for 70B Llama model.[3]

🔮 前景展望AI analysis grounded in cited sources

Together GPU Clusters will capture more enterprise AI workloads from static provisioning providers.
Dynamic autoscaling and self-healing address key pain points of over/underprovisioning and node failures, enabling efficient scaling for variable training and inference.[1][2]
$1B funding pursuit will accelerate Together AI's GPU infrastructure expansion.
The upgrades coincide with reported funding talks, providing production-ready features to attract enterprise customers and investors in the competitive GPU-as-a-service market.[1][7]

時間線

2025-03
Introduced Instant GPU Clusters at NVIDIA GTC 2025.
2025-09
Launched self-service GPU infrastructure with Together Instant Clusters.
2026-03
Announced autoscaling, RBAC, observability, and self-healing for GPU Clusters.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Together AI Blog

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。