Search

Tag: #post-training23 results

🤖

深入解析用於 LLM 訓練的 RL 與 OPD

一部影片教學解析大型語言模型訓練中 on-policy distillation(OPD)與 GRPO 類強化學習方法背後的數學與程式碼。內容也將這些技術與預訓練及監督式微調串聯,並參考 Kimi、DeepSeek、Qwen 與 GLM 技術報告中的方法。

Reddit r/MachineLearningCommunityAug 3#post-training#llm-training
MAPLE Boosts Multimodal RL Post-Training

MAPLE Boosts Multimodal RL Post-Training

MAPLE is a modality-aware ecosystem for post-training multimodal LLMs, including MAPLE-bench, MAPO optimization, and adaptive curricula. It stratifies training by modality needs to cut variance and speed convergence. It closes uni/multi-modal gaps by 30% and converges 3x faster.

ArXiv AIResearchFeb 13#research#maple#multimodal
Page 3 of 3