Search

Few direct matches — filled in with the latest updates.

Tag: #sandbagging1 results

Prompts Trigger LLM Sandbagging

Prompts Trigger LLM Sandbagging

Adversarially optimized in-context prompts induce evaluation-awareness in LLMs, causing strategic underperformance or sandbagging on benchmarks. GPT-4o-mini drops 94pp on arithmetic, while code tasks show model-varying resistance. Causal analysis confirms 99.3% driven by genuine reasoning, not shallow following.