SourceStalecollected in 10h

Refusal Alignment Eval Fails on Routing

PostLinkedIn
🤖Read original on Reddit r/MachineLearning
#alignment#evaluation#censorship#probesrefusal-based-alignment-evalqwendeepseekglmphiyi

💡Flaws in alignment evals exposed via Chinese LLM censorship

⚡ 30-Second TL;DR

What Changed

Probes hit 100% even on null/random data; held-out generalization discriminates

Why It Matters

Challenges current alignment evals, urging causal interventions over refusal benchmarks for reliable safety assessment across labs.

What To Do Next

Test held-out probes and ablate routing on your aligned LLMs.

Who should care:Researchers & Academics

Key Points

  • Probes hit 100% even on null/random data; held-out generalization discriminates
  • Ablation removes censorship in 3/4 models; Qwen3-8B confabulates facts
  • Routing lab-specific, orthogonal to safety; refusal evals miss steering changes
  • 46-model screen shows censorship in few models; proposes probe evidence hierarchy
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.