Chain-of-Thought Control Tests May Underestimate Models

Better prompts may double or triple measured CoT control, challenging current model-safety conclusions.
30-Second TL;DR
What Changed
Recent models reportedly score between 0% and 30% on CoTControl, with one exception.
Why It Matters
Safety conclusions based on low CoT controllability scores may be too optimistic if evaluations fail to elicit model capabilities. More robust testing could change how labs assess monitorability and reasoning-trace oversight.
What To Do Next
Re-run your CoT monitoring evaluations with optimized prompt templates and training-based elicitation before drawing safety conclusions.
Key Points
- •Recent models reportedly score between 0% and 30% on CoTControl, with one exception.
- •Prompt iteration improved some open-source model scores by approximately two to three times.
- •The analysis recommends stronger prompt optimization and training-based elicitation.
Deep Insight
Background and context from public sources — not the original article. 10 sources cited.
Enhanced Key Takeaways
- •The formal CoT-Control evaluation suite was introduced by Yueh-Han Chen et al. at ICML 2026 to measure whether reasoning models can adhere to explicit scratchpad constraints, such as solving problems while avoiding forbidden tokens.
- •The under-elicitation discovery showed that automated prompt optimization via Claude Opus 4.6 lifted open-weight models like GPT-OSS-120B from a 5.5% baseline to 15% constraint compliance.
- •Frontier models demonstrate a severe disparity between scratchpad and output control; Claude Sonnet 4.5 achieved only 2.7% compliance within its CoT despite reaching over 60% compliance on identical constraints in final visible outputs.
- •The sole noted exception to the 0% to 30% CoT controllability ceiling was Claude Mythos Preview, which demonstrated elevated steerability after accidental CoT exposure during reinforcement learning training.
- •Frontier system filings for models like GPT-5.5 and Fable 5 have cited low CoT-Control scores to argue models cannot evade safety monitors, despite the UK AI Safety Institute noting extreme benchmark sensitivity to prompt scaffolding.
Technical Deep Dive
- Benchmark Structure: CoT-Control evaluates negative and structural constraints in model scratchpads, requiring models to solve multi-step problems (e.g., genetics queries) without emitting specific required domain terms (e.g., 'chromosome').
- Meta-Cognitive Failure Mode: Baseline models frequently demonstrate awareness of constraints within their reasoning trace (explicitly generating phrases like 'Wait, I am not supposed to use that word') yet proceed to generate the forbidden tokens due to token-level predictive momentum needed for problem resolution.
- Elicitation Discrepancy: Models routinely show >60% compliance when enforcing constraints on visible final generation, but plummet to 2.7%–15.4% inside unmonitored or intermediate reasoning chains.
- Optimization Methodology: Replacing raw, zero-shot instructions with meta-prompted scaffolds generated by advanced models (Claude Opus 4.6) doubled or tripled constraint adherence without architectural modifications or weight updates.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-03OpenAI publishes research highlighting that frontier reasoning models struggle to control scratchpads
- 2026-06Yueh-Han Chen et al. publish formal CoT-Control benchmark suite at ICML 2026
- 2026-08Fable 5 system card discloses UK AISI findings on CoT prompt scaffolding sensitivity
- 2026-09AI Alignment Forum analysis reveals CoTControl is under-elicited and improves 2-3x via prompt optimization
Sources (10)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: AI Alignment Forum ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.