Rethinking Reasoning SFT Generalization

💡SFT can generalize like RL—under right optimization, data, models. Key for fine-tuners.
⚡ 30-Second TL;DR
What Changed
Cross-domain performance dips then recovers with longer SFT training
Why It Matters
Reframing SFT for reasoning prompts practitioners to prioritize data quality and extended training. Highlights trade-offs like safety degradation, influencing LLM deployment strategies.
What To Do Next
Extend SFT training epochs on reasoning datasets to observe dip-and-recovery generalization gains.
Key Points
- •Cross-domain performance dips then recovers with longer SFT training
- •Verified long-CoT data boosts generalization across domains
- •Stronger models learn transferable patterns like backtracking from toy tasks
- •Generalization asymmetric: improves reasoning but degrades safety
🧠 Deep Insight
AI-generated analysis for this event — not the original article.
🔑 Enhanced Key Takeaways
- •The 'dip-and-recovery' phenomenon is linked to the interference between pre-trained knowledge and new reasoning trajectories, where initial SFT steps cause catastrophic forgetting of general capabilities before the model learns to integrate the new CoT format.
- •Research indicates that the degradation of safety during reasoning SFT is primarily due to the model prioritizing the 'reasoning' objective function over the 'harmlessness' constraints embedded during RLHF, suggesting a need for multi-objective fine-tuning.
- •The effectiveness of cross-domain generalization is highly sensitive to the 'reasoning density' of the training data; models trained on sparse reasoning chains fail to generalize, whereas dense, multi-step chains facilitate the emergence of transferable heuristic search strategies.
🛠️ Technical Deep Dive
- •Training dynamics analysis shows that the 'dip' phase corresponds to a spike in loss on non-reasoning benchmarks, suggesting a temporary collapse of the model's internal representation space.
- •The recovery phase is characterized by the alignment of the model's hidden states with the structure of the new reasoning tasks, effectively 're-mapping' pre-trained knowledge into the new CoT format.
- •Asymmetric safety degradation is quantified by a significant increase in jailbreak success rates when models are fine-tuned on reasoning tasks without concurrent safety-alignment regularization.
- •Backtracking capability emergence is correlated with the depth of the CoT chains in the training set, with a threshold effect observed at approximately 500-1000 reasoning steps per sample.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.