🤖Stalecollected in 5h

RLVR Beats SFT on Qwen2.5 Math +11.9pts

RLVR Beats SFT on Qwen2.5 Math +11.9pts
PostLinkedIn
🤖Read original on Reddit r/MachineLearning
#fine-tuning#reasoning#rlhfqwen2.5-1.5b-rlvrqwen2.5-1.5brlvrgrpogsm8ksft

💡RLVR +12pts math on Qwen2.5 from 1 example—SFT hurts reasoning. Checkpoints live.

⚡ 30-Second TL;DR

What Changed

RLVR +11.9pts GSM8K, generalizes to MATH even from 1 example

Why It Matters

Demonstrates RLVR's superiority for reasoning over SFT, enabling efficient training with minimal data; influences fine-tuning strategies for open models.

What To Do Next

Fine-tune your model with GRPO from github.com/jayminban/RLVR-vs-SFT-Qwen2.5-1.5b.

Who should care:Researchers & Academics

Key Points

  • RLVR +11.9pts GSM8K, generalizes to MATH even from 1 example
  • SFT -15.2pts GSM8K, overrides pretrained knowledge
  • 388 checkpoints benchmarked, 2.4M rows in SQLite DB
  • GRPO used on 6-8x RTX GPUs, code/checkpoints on GitHub
  • Reduces no-answer rate but harms accuracy

🧠 Deep Insight

Background and context from public sources — not the original article. 8 sources cited.

🔑 Enhanced Key Takeaways

  • Qwen2.5 series pretrained on up to 18 trillion tokens, achieving MMLU scores over 85, HumanEval over 85, and MATH over 80, marking substantial knowledge and capability gains over Qwen2[3][4].
  • Qwen2.5-Math-72B-Instruct surpasses Qwen2-Math-72B-Instruct and GPT-4o on general performance, with even the 1.5B-Instruct variant competitive against much larger models[3].
  • Qwen2.5 enhances post-training for 8K token generation, structured data comprehension like tables, reliable JSON outputs, and resilience to diverse system prompts for role-playing[3][4].

🔮 Future ImplicationsAI analysis grounded in cited sources

RLVR will become standard for fine-tuning small math-specialized models under 3B parameters.
The +11.9pt gain on Qwen2.5-1.5B with minimal examples demonstrates RLVR's sample efficiency over SFT, enabling accessible deployment on consumer hardware.
GRPO-based methods will reduce catastrophic forgetting in instruction tuning by 2026 Q3.
SFT's -15.2pt degradation highlights overriding of pretrained knowledge, positioning RLVR as a solution for preserving base capabilities during specialization.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.