🧠机器之心•較早收集於 7m
Co-rewarding:無標註穩定RL誘導LLM推理

#self-supervised-rl#reward-hacking#llm-reasoningco-rewarding
💡ICLR 2026 breakthrough: stable label-free RL for LLMs, beats reward hacking.
⚡ 30-Second TL;DR
有什麼變化
ICLR 2026接收Co-rewarding自監督RL框架
為什麼重要
實現無昂貴標註的RL訓練擴展推理,潛在加速研究人員與開發者面對數據短缺的LLM進展。
下一步行動
Read the paper at https://openreview.net/forum?id=fDk95XPsCU and test Co-rewarding on your LLM RL pipeline.
誰應關注:Researchers & Academics
關鍵要點
- •ICLR 2026接收Co-rewarding自監督RL框架
- •透過數據/模型視角互補訊號消除RLVR標註需求
- •防止自獎勵LLM的訓練崩潰與獎勵駭客
- •誘導大語言模型穩定推理能力
- •提供論文與程式碼連結供複現
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 6 個來源。
🔑 增強重點摘要
- •Co-rewarding is a self-supervised RL framework for LLM reasoning, accepted to ICLR 2026, developed by researchers from Hong Kong Baptist University and Shanghai Jiao Tong University.
- •It addresses RLVR annotation bottlenecks and self-rewarding failures by using complementary self-supervised signals from data and model views, eliminating the need for ground-truth labels.
- •The framework prevents reward hacking and training collapse, inducing stable reasoning capabilities in large language models.
- •ICLR 2026 features 5344 accepted papers from 18949 submissions (28.20% acceptance rate), highlighting competitive selection for works like Co-rewarding[1].
- •Related ICLR 2026 papers tackle similar reward hacking issues in RL for LLMs and diffusion models, indicating active research in stable RL paradigms[2][4][5].
📊 競品分析▸ Show
| Method | Key Feature | Benchmarks |
|---|---|---|
| Co-rewarding | Label-free self-supervised RL with complementary signals | Prevents collapse in LLM reasoning (no specific benchmarks in results) |
| REA-RL | Reflection-aware online RL for efficient reasoning models | ICLR 2026 acceptance [2] |
| SMORM | Joint Bradley-Terry and multi-objective reward modeling | Outperforms 70B baseline with 7B model in OOD settings [4] |
| IntDiff | Intrinsic rewards for diffusion model fine-tuning | Improves alignment and diversity in text-to-image [5] |
| LongR | Contextual dense rewards for long-context reasoning | ~4% gain over outcome-only baselines [3] |
🛠️ 技術深入
- •Co-rewarding employs complementary self-supervised signals from data/model perspectives to stabilize RL training without labels, targeting reward hacking prevention in LLM reasoning induction.
- •No specific model architecture or implementation details (e.g., code links) found in search results beyond original article mention.
- •Related works: LongR uses Relative Information Gain (white-box metric) and interleaved Think-and-Read policy with curriculum learning for long-context RLVR[3].
- •REA-RL provides GitHub implementation for reflection-aware online RL, accepted to ICLR 2026[2].
🔮 前景展望AI analysis grounded in cited sources
Co-rewarding advances label-free RL for LLMs, potentially reducing annotation costs and improving reasoning stability amid growing ICLR focus on reward modeling robustness, enabling scalable self-improvement in reasoning models.
📎 來源 (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心 ↗
每週 AI 簡報
每週一封,可隨時退訂。