SourceStalecollected in 3h

LMs' Blind Refusal to Unjust Rules

LMs' Blind Refusal to Unjust Rules
PostLinkedIn
📄Read original on ArXiv AI
#blind-refusal#lm-safety#normative-reasoning#alignmentllmsarxivgpt-5.4

💡LMs refuse 75% justified rule evasions—key alignment flaw exposed

⚡ 30-Second TL;DR

What Changed

Dataset crosses 5 rule defeat families with 19 authority types

Why It Matters

Highlights decoupling of normative reasoning from behavior in LMs, challenging current safety training. Informs alignment research to enable justified non-compliance without risking misuse.

What To Do Next

Download arXiv:2404.06233 dataset to benchmark your LM's blind refusal.

Who should care:Researchers & Academics

Key Points

  • Dataset crosses 5 rule defeat families with 19 authority types
  • Tested 18 model configs from 7 families, N=14,650 responses
  • 75.4% refusal rate on defeated-rule requests without safety risks
  • 57.5% models recognize defeat conditions but still decline help
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.