🤝Together AI Blog•較早收集於 34h
Together GPU Clusters 新增自動擴展與自癒
.png)
#autoscaling#observability#self-healing#rbactogether-gpu-clusterstogether-ai
💡具自動擴展與自癒的生產 GPU 叢集,可降低 AI 基礎設施成本 30-50%。
⚡ 30-Second TL;DR
有什麼變化
內建自動擴展實現動態資源調整
為什麼重要
這降低 AI 團隊的運維負擔,讓他們專注開發而非維護。提升大規模 ML 訓練與推論的成本效益與可靠性。
下一步行動
註冊 Together GPU Clusters,並在範例 ML 工作負載上測試自動擴展。
誰應關注:Enterprise & Security Teams
關鍵要點
- •內建自動擴展實現動態資源調整
- •RBAC 提供安全角色存取控制
- •全棧觀測性實現全面監控
- •自癒節點修復自動故障恢復
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 8 個來源。
🔑 增強重點摘要
- •Together AI is reportedly pursuing a $1 billion funding round amid these infrastructure upgrades.[1]
- •Autoscaling is powered by Kubernetes Cluster Autoscaler, which monitors GPU-constrained workloads and automatically provisions or decommissions nodes based on real-time demand.[1][2]
- •Self-healing involves a three-click process where the control plane cords, drains, and recreates failed nodes, with automatic acceptance tests ensuring clusters pass before marking as ready.[2]
- •GPU Clusters support NVIDIA H100, H200, B200, and GB200 GPUs with InfiniBand networking, scaling from 8 to over 4,000 GPUs for distributed training and inference.[3]
🛠️ 技術深入
- •Autoscaling uses Kubernetes Cluster Autoscaler for dynamic node provisioning/deprovisioning based on GPU demand, targeting variable inference and bursty training jobs.[1][2]
- •Observability features a dedicated Grafana instance with dashboards for DCGM GPU metrics, InfiniBand/NIC networking telemetry, storage I/O, and Kubernetes health.[2]
- •Self-healing node repair: Cordon and drain failed nodes, recreate on new/existing hosts, run acceptance tests (detailed in documentation) before readiness.[2]
- •Orchestration options include managed Kubernetes (Kubeadm-based, HA control plane, node autoscaling, Cert Manager) and Slurm on Kubernetes for precise gang scheduling and SSH access.[3]
- •Performance enhanced by Together Kernel Collection (including ThunderKittens), achieving 90% faster training on B200 (15,264 tokens/s/GPU) vs H100 (8,080 tokens/s/GPU) for 70B Llama model.[3]
🔮 前景展望AI analysis grounded in cited sources
Together GPU Clusters will capture more enterprise AI workloads from static provisioning providers.
⏳ 時間線
2025-03
Introduced Instant GPU Clusters at NVIDIA GTC 2025.
2025-09
Launched self-service GPU infrastructure with Together Instant Clusters.
2026-03
Announced autoscaling, RBAC, observability, and self-healing for GPU Clusters.
📎 來源 (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- mexc.com — 897572
- together.ai — New in Together GPU Clusters Autoscaling Observability Self Healing
- together.ai — GPU Clusters
- together.ai
- together.ai — AI Native Conf Research and Product Announcements
- prnewswire.com — Together AI Announces Business and Product Milestones at First AI Native Conference 302705505
- datacenterdynamics.com — Together AI Seeks 1bn in Funding Report
- together.ai — Nvidia Gtc 2026
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Together AI Blog ↗
每週 AI 簡報
每週一封,可隨時退訂。