來源量子位•較早收集於 40m
商湯大裝置重塑 AI 算力叢集

#ai-native#compute-clusters#cloud-infrasensetime-big-devicesensetime
商湯 AI 原生時代叢集重構—擴展運算基礎設施關鍵。
30 秒速覽
有什麼變化
推出 AI 原生算力叢集重新設計
為什麼重要
實現大規模 AI 訓練更高效,降低 AI 企業超大規模運算成本。
下一步行動
探索商湯 AI 原生雲文件,優化基礎設施堆疊叢集。
誰應關注:Enterprise & Security Teams
關鍵要點
- •推出 AI 原生算力叢集重新設計
- •商湯大裝置引領架構革新
- •分享 AI 原生雲部署實踐
深度解析
本篇為 AI 生成分析,非原文內容。
增強重點摘要
- •SenseTime's 'SenseCore' AI infrastructure platform serves as the foundational layer for the Big Device, integrating massive-scale GPU resource scheduling with high-performance storage and networking to support training models exceeding 1 trillion parameters.
- •The architecture emphasizes 'AI-native' design by optimizing the interaction between the compute layer and the data layer, specifically addressing the bottleneck of data throughput during large-scale distributed training of multimodal foundation models.
- •The Big Device utilizes a proprietary high-speed interconnect fabric that significantly reduces latency in collective communication operations (like AllReduce) compared to standard off-the-shelf networking solutions, enabling higher GPU utilization rates.
競品分析
Primary Focus
- SenseTime Big Device
- AI-native cloud/model training
- NVIDIA DGX SuperPOD
- Turnkey enterprise AI infrastructure
- Huawei Ascend AI Cluster
- Domestic compute sovereignty/Ascend chips
Interconnect
- SenseTime Big Device
- Proprietary high-speed fabric
- NVIDIA DGX SuperPOD
- NVLink / InfiniBand
- Huawei Ascend AI Cluster
- HCCS / RoCE
Software Stack
- SenseTime Big Device
- SenseCore
- NVIDIA DGX SuperPOD
- NVIDIA AI Enterprise / Base Command
- Huawei Ascend AI Cluster
- CANN / MindSpore
| Feature | SenseTime Big Device | NVIDIA DGX SuperPOD | Huawei Ascend AI Cluster |
|---|---|---|---|
| Primary Focus | AI-native cloud/model training | Turnkey enterprise AI infrastructure | Domestic compute sovereignty/Ascend chips |
| Interconnect | Proprietary high-speed fabric | NVLink / InfiniBand | HCCS / RoCE |
| Software Stack | SenseCore | NVIDIA AI Enterprise / Base Command | CANN / MindSpore |
技術深入
- Compute Density: Optimized for high-density GPU clusters, supporting multi-thousand GPU nodes in a single training job.
- Data Throughput: Implements a tiered storage architecture that separates hot/cold data to minimize I/O wait times during checkpointing and model loading.
- Scheduling: Features a custom-built scheduler designed to handle heterogeneous workloads, allowing for dynamic resource allocation between model training and inference tasks.
- Communication: Utilizes advanced topology-aware routing to minimize network congestion in large-scale distributed training environments.
前景展望基於引用來源的 AI 分析
SenseTime will transition toward a model-as-a-service (MaaS) dominant revenue model.
The efficiency gains from the Big Device infrastructure lower the cost of training and serving proprietary foundation models, making MaaS more economically viable.
The Big Device architecture will become the standard for domestic Chinese AI cloud providers.
As access to high-end Western networking hardware remains constrained, SenseTime's proprietary interconnect and cluster management software offer a critical alternative for scaling AI compute.
時間線
2022-01
SenseTime officially launches SenseCore AI infrastructure platform.
2023-04
SenseTime unveils 'SenseNova' foundation model suite, necessitating the scaling of Big Device infrastructure.
2024-07
SenseTime announces significant upgrades to its AI computing cluster capacity to support 100B+ parameter model training.
- 2022-01SenseTime officially launches SenseCore AI infrastructure platform.
- 2023-04SenseTime unveils 'SenseNova' foundation model suite, necessitating the scaling of Big Device infrastructure.
- 2024-07SenseTime announces significant upgrades to its AI computing cluster capacity to support 100B+ parameter model training.
AI 週報
閱讀本週精選 AI 大事摘要 →
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 量子位 ↗
每週電子報
每週一封,可隨時退訂。