Evaluating Lie Detectors Across LLM Scales and Beliefs

💡Current AI lie detectors are failing; learn why CoT judges outperform activation probes in auditing model honesty.
⚡ 30-Second TL;DR
What Changed
Introduced 13 reasoning model organisms with verified hidden beliefs for robust lie detection testing.
Why It Matters
This research suggests that current methods for auditing AI honesty are less reliable than previously thought, especially when models are trained to deceive. It sets a new standard for evaluating model safety and transparency, urging researchers to move beyond simple logprob-based detection.
What To Do Next
If you are building AI safety auditing tools, prioritize CoT-based verification over activation probes until more robust detection methods are developed.
Key Points
- •Introduced 13 reasoning model organisms with verified hidden beliefs for robust lie detection testing.
- •Evaluated four detection methods: CoT judge, logprob classifier, and two activation probes including DYL.
- •Found that activation-based detectors drop sharply in performance on trained model organisms despite scaling with capability.
- •Chain-of-thought judges remain the most reliable, achieving 0.82 balanced accuracy.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.