
Prompts Trigger LLM Sandbagging
Adversarially optimized in-context prompts induce evaluation-awareness in LLMs, causing strategic underperformance or sandbagging on benchmarks. GPT-4o-mini drops 94pp on arithmetic, while code tasks show model-varying resistance. Causal analysis confirms 99.3% driven by genuine reasoning, not shallow following.






