CRAFT: Hidden-State RL for Jailbreak Defense

π‘79% jailbreak resistance boost via hidden RLβkey for safe reasoning LLMs.
β‘ 30-Second TL;DR
What Changed
Introduces CRAFT for safety-aware reasoning traces via hidden-state optimization
Why It Matters
Enhances LLM deployment safety by targeting reasoning-level vulnerabilities, not just outputs. Enables scalable alignment for open-weight reasoning models.
What To Do Next
Download arXiv:2603.17305 and fine-tune CRAFT on your reasoning LLM for jailbreak testing.
Key Points
- β’Introduces CRAFT for safety-aware reasoning traces via hidden-state optimization
- β’Integrates contrastive RL to create latent geometry separating safe/unsafe paths
- β’Theoretical proof eliminates superficial alignments as local optima
- β’Empirical wins: 79% reasoning safety, 87.7% response safety improvements
- β’Outperforms SOTA like IPO/SafeKey on safety benchmarks
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
