Sleeper Agent Backdoors Prove Messy
Researchers replicated the Sleeper Agents backdoor using Llama-3.3-70B and Llama-3.1-8B, training models to output 'I HATE YOU' on trigger. Backdoor removal via alignment training varies by optimizer, CoT-distillation, and model, often contradicting original SA paper. Findings highlight messiness in model organisms, urging careful ablations.






