SourceStalecollected in 21h

Rethinking Reasoning SFT Generalization

Rethinking Reasoning SFT Generalization
PostLinkedIn
📄Read original on ArXiv AI
#generalization#chain-of-thoughtllm-sftarxivllmsftcot

💡SFT can generalize like RL—under right optimization, data, models. Key for fine-tuners.

⚡ 30-Second TL;DR

What Changed

Cross-domain performance dips then recovers with longer SFT training

Why It Matters

Reframing SFT for reasoning prompts practitioners to prioritize data quality and extended training. Highlights trade-offs like safety degradation, influencing LLM deployment strategies.

What To Do Next

Extend SFT training epochs on reasoning datasets to observe dip-and-recovery generalization gains.

Who should care:Researchers & Academics

Key Points

  • Cross-domain performance dips then recovers with longer SFT training
  • Verified long-CoT data boosts generalization across domains
  • Stronger models learn transferable patterns like backtracking from toy tasks
  • Generalization asymmetric: improves reasoning but degrades safety

🧠 Deep Insight

AI-generated analysis for this event — not the original article.

🔑 Enhanced Key Takeaways

  • The 'dip-and-recovery' phenomenon is linked to the interference between pre-trained knowledge and new reasoning trajectories, where initial SFT steps cause catastrophic forgetting of general capabilities before the model learns to integrate the new CoT format.
  • Research indicates that the degradation of safety during reasoning SFT is primarily due to the model prioritizing the 'reasoning' objective function over the 'harmlessness' constraints embedded during RLHF, suggesting a need for multi-objective fine-tuning.
  • The effectiveness of cross-domain generalization is highly sensitive to the 'reasoning density' of the training data; models trained on sparse reasoning chains fail to generalize, whereas dense, multi-step chains facilitate the emergence of transferable heuristic search strategies.

🛠️ Technical Deep Dive

  • Training dynamics analysis shows that the 'dip' phase corresponds to a spike in loss on non-reasoning benchmarks, suggesting a temporary collapse of the model's internal representation space.
  • The recovery phase is characterized by the alignment of the model's hidden states with the structure of the new reasoning tasks, effectively 're-mapping' pre-trained knowledge into the new CoT format.
  • Asymmetric safety degradation is quantified by a significant increase in jailbreak success rates when models are fine-tuned on reasoning tasks without concurrent safety-alignment regularization.
  • Backtracking capability emergence is correlated with the depth of the CoT chains in the training set, with a threshold effect observed at approximately 500-1000 reasoning steps per sample.

🔮 Future ImplicationsAI analysis grounded in cited sources

Future SFT protocols will mandate concurrent safety-alignment to prevent reasoning-induced degradation.
The observed asymmetric safety loss necessitates a shift from sequential to multi-objective training pipelines.
Reasoning-focused datasets will shift toward 'dense-CoT' formats to maximize cross-domain transfer.
Empirical evidence shows that dense reasoning chains are required for the model to learn transferable heuristic search patterns.

Timeline

2024-09
Initial research into reasoning-based SFT reveals the 'dip' phenomenon in cross-domain benchmarks.
2025-03
Introduction of multi-objective SFT frameworks to mitigate safety degradation during reasoning training.
2025-11
Publication of findings linking reasoning density in training data to the emergence of backtracking capabilities.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.