Hard CoT Interp Tasks Released

💡9 OOD-hard CoT tasks: probes beat LLMs—benchmark your interp tools!
⚡ 30-Second TL;DR
What Changed
9 tasks for CoT interp: predict stopping, sycophancy, confidence, etc.
Why It Matters
Provides standardized OOD benchmark for CoT tools, vital for AI safety techniques. Highlights probes/TF-IDF as strong baselines, spurring non-LLM method development.
What To Do Next
Download datasets from the AI Alignment Forum repo and baseline your CoT interp method on the 7 main OOD tasks.
Key Points
- •9 tasks for CoT interp: predict stopping, sycophancy, confidence, etc.
- •Open-sourced datasets/code with in-dist/OOD splits.
- •Baselines: TF-IDF/attention probes beat zero/few-shot LLM monitors OOD.
- •7 main + 2 misc tasks to prove OOD robustness.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: AI Alignment Forum ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.