Elicitation vs Creation in LLM Post-Training

π‘New framework to check if post-training elicits or creates LLM capabilities.
β‘ 30-Second TL;DR
What Changed
Introduces 'accessible support': behaviors reachable under finite compute budgets.
Why It Matters
Clarifies SFT vs RL debates, enabling better evaluation of true capability gains. Guides choice of post-training methods for genuine expansion vs optimization.
What To Do Next
Read arXiv:2605.08368 and test accessible support in your next SFT/RL experiment.
Key Points
- β’Introduces 'accessible support': behaviors reachable under finite compute budgets.
- β’Elicitation reweights within support; creation alters the support itself.
- β’SFT/RL both reweight pretrained distribution using demo/reward signals.
- β’Local post-training near base model mainly elicits, not creates.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.