Search

Tag: #distributed-training23 results

DeepSeek 更新 DeepGEMM 支援 Mega MoE

DeepSeek 更新 DeepGEMM 支援 Mega MoE

DeepSeek 更新其 DeepGEMM 儲存庫,新增對 Mega MoE 的測試支援,包括 P4 量化、分佈式通訊、Blackwell 適配以及 HyperConnection 訓練。這暗示正準備部署比 V3 更大的 MoE 模型,可能為 DeepSeek V4,需要 FP4 以實現高效推論。此更新僅限 DeepGEMM 開發,與內部模型發布無關。

Reddit r/LocalLLaMACommunityApr 16#moe#fp4#blackwell
使用 veRL 和 Ray 在 SageMaker 訓練 CodeFu-7B

使用 veRL 和 Ray 在 SageMaker 訓練 CodeFu-7B

這篇文章展示如何使用 veRL 的 GRPO 在 SageMaker 管理的 Ray 分佈式叢集上訓練 CodeFu-7B,一個專為競賽程式設計的 70 億參數模型。它涵蓋資料準備、分佈式訓練設定和全面觀測性。此統一方法為複雜 RL 訓練工作負載提供計算規模和開發者體驗。

AWS Machine Learning BlogOfficialFeb 24#distributed-training#rl-optimization
Page 2 of 3