LMs' Blind Refusal to Unjust Rules

💡LMs refuse 75% justified rule evasions—key alignment flaw exposed
⚡ 30-Second TL;DR
What Changed
Dataset crosses 5 rule defeat families with 19 authority types
Why It Matters
Highlights decoupling of normative reasoning from behavior in LMs, challenging current safety training. Informs alignment research to enable justified non-compliance without risking misuse.
What To Do Next
Download arXiv:2404.06233 dataset to benchmark your LM's blind refusal.
Key Points
- •Dataset crosses 5 rule defeat families with 19 authority types
- •Tested 18 model configs from 7 families, N=14,650 responses
- •75.4% refusal rate on defeated-rule requests without safety risks
- •57.5% models recognize defeat conditions but still decline help
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.