Search

Tag: #rlvr5 results

Multilingual Reasoning Gym for 14 Languages

Multilingual Reasoning Gym for 14 Languages

Apple's Multilingual Reasoning Gym extends the original Reasoning Gym to generate verifiable reasoning problems across 14 languages. It features translated templates for 94 tasks, validated by native speakers in 10 languages with adaptations for naturalness. The gym retains procedural generation for unlimited instances and adjustable difficulty, ideal for reinforcement learning.

Apple Machine LearningOfficialMar 13#multilingual#reasoning#benchmark
Multilingual Math Dataset for RLVR

Multilingual Math Dataset for RLVR

Apple's mAceReason-Math provides high-quality multilingual math problems designed for Reinforcement Learning with Verifiable Rewards (RLVR). It addresses the English-centric bias in existing datasets, offering appropriate difficulty for current LLMs. The dataset supports boosting math and logic capabilities in pretrained models.

Apple Machine LearningOfficialMar 13#multilingual#math#dataset
🔬

WizardLM Releases Mix-GRM Paper

WizardLM released a new paper on improving Generative Reward Models by synergizing Breadth (B-CoT) and Depth (D-CoT) reasoning instead of just longer CoT. The Mix-GRM framework uses RLVR training to enable the model to autonomously select reasoning structures, achieving high performance with efficient token use. It addresses limitations in subjective vs. objective evaluation tasks.

Reddit r/LocalLLaMACommunityMar 4#reward-models#cot-reasoning#rlvr
GradLoc Locates RLVR Crash Tokens

GradLoc Locates RLVR Crash Tokens

Tencent Hunyuan's GradLoc tool pinpoints gradient spikes to specific tokens in RLVR training, ending guesswork for collapses. Provides infrastructure for observable RL dynamics in high-noise systems. Open-sourced GitHub lowers engineering barriers for mechanism research.

机器之心MediaFeb 14#research#hunyuan#gradloc