SourceReddit r/MachineLearning•Stalecollected in 10h
Refusal Alignment Eval Fails on Routing
#alignment#evaluation#censorship#probesrefusal-based-alignment-evalqwendeepseekglmphiyi
💡Flaws in alignment evals exposed via Chinese LLM censorship
⚡ 30-Second TL;DR
What Changed
Probes hit 100% even on null/random data; held-out generalization discriminates
Why It Matters
Challenges current alignment evals, urging causal interventions over refusal benchmarks for reliable safety assessment across labs.
What To Do Next
Test held-out probes and ablate routing on your aligned LLMs.
Who should care:Researchers & Academics
Key Points
- •Probes hit 100% even on null/random data; held-out generalization discriminates
- •Ablation removes censorship in 3/4 models; Qwen3-8B confabulates facts
- •Routing lab-specific, orthogonal to safety; refusal evals miss steering changes
- •46-model screen shows censorship in few models; proposes probe evidence hierarchy
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.