
Anchored Bipolicy Self-Play Boosts AI Safety
Researchers propose Anchored Bipolicy Self-Play to address limitations in standard self-play red teaming for AI safety. It uses distinct role-specific LoRA adapters on a frozen base model, avoiding self-consistency collapse and maintaining adversarial pressure. Evaluations on Qwen2.5 models show superior safety and efficiency over baselines.