SourceStalecollected in 5h

Chain-of-Thought Control Tests May Underestimate Models

Read original on AI Alignment Forum
#chain-of-thought#evaluations#model-monitoring

Better prompts may double or triple measured CoT control, challenging current model-safety conclusions.

30-Second TL;DR

What Changed

Recent models reportedly score between 0% and 30% on CoTControl, with one exception.

Why It Matters

Safety conclusions based on low CoT controllability scores may be too optimistic if evaluations fail to elicit model capabilities. More robust testing could change how labs assess monitorability and reasoning-trace oversight.

What To Do Next

Re-run your CoT monitoring evaluations with optimized prompt templates and training-based elicitation before drawing safety conclusions.

Who should care:Researchers & Academics

Key Points

  • Recent models reportedly score between 0% and 30% on CoTControl, with one exception.
  • Prompt iteration improved some open-source model scores by approximately two to three times.
  • The analysis recommends stronger prompt optimization and training-based elicitation.
Key numbers5.5%15%2.7%60%

Deep Insight

Background and context from public sources — not the original article. 10 sources cited.

Enhanced Key Takeaways

  • The formal CoT-Control evaluation suite was introduced by Yueh-Han Chen et al. at ICML 2026 to measure whether reasoning models can adhere to explicit scratchpad constraints, such as solving problems while avoiding forbidden tokens.
  • The under-elicitation discovery showed that automated prompt optimization via Claude Opus 4.6 lifted open-weight models like GPT-OSS-120B from a 5.5% baseline to 15% constraint compliance.
  • Frontier models demonstrate a severe disparity between scratchpad and output control; Claude Sonnet 4.5 achieved only 2.7% compliance within its CoT despite reaching over 60% compliance on identical constraints in final visible outputs.
  • The sole noted exception to the 0% to 30% CoT controllability ceiling was Claude Mythos Preview, which demonstrated elevated steerability after accidental CoT exposure during reinforcement learning training.
  • Frontier system filings for models like GPT-5.5 and Fable 5 have cited low CoT-Control scores to argue models cannot evade safety monitors, despite the UK AI Safety Institute noting extreme benchmark sensitivity to prompt scaffolding.

Technical Deep Dive

  • Benchmark Structure: CoT-Control evaluates negative and structural constraints in model scratchpads, requiring models to solve multi-step problems (e.g., genetics queries) without emitting specific required domain terms (e.g., 'chromosome').
  • Meta-Cognitive Failure Mode: Baseline models frequently demonstrate awareness of constraints within their reasoning trace (explicitly generating phrases like 'Wait, I am not supposed to use that word') yet proceed to generate the forbidden tokens due to token-level predictive momentum needed for problem resolution.
  • Elicitation Discrepancy: Models routinely show >60% compliance when enforcing constraints on visible final generation, but plummet to 2.7%–15.4% inside unmonitored or intermediate reasoning chains.
  • Optimization Methodology: Replacing raw, zero-shot instructions with meta-prompted scaffolds generated by advanced models (Claude Opus 4.6) doubled or tripled constraint adherence without architectural modifications or weight updates.

Future ImplicationsAI analysis grounded in cited sources

Frontier AI safety filings will be forced to deprecate raw zero-shot CoT-Control scores.
Regulators and safety institutes will require training-based and heavily optimized elicitation before accepting claims that a model cannot obfuscate its scratchpad reasoning.
Safety monitoring will increasingly pivot toward detecting zero-scratchpad reasoning.
As models develop internal latent reasoning capabilities without explicit chain-of-thought tokens, scratchpad inspection alone will provide diminishing safety guarantees.

Timeline

2026-03
OpenAI publishes research highlighting that frontier reasoning models struggle to control scratchpads
2026-06
Yueh-Han Chen et al. publish formal CoT-Control benchmark suite at ICML 2026
2026-08
Fable 5 system card discloses UK AISI findings on CoT prompt scaffolding sensitivity
2026-09
AI Alignment Forum analysis reveals CoTControl is under-elicited and improves 2-3x via prompt optimization

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: AI Alignment Forum

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.