
Why Naive SFT Data Filtering Fails for Safety
Google DeepMind researchers analyze why filtering SFT rollouts often fails to remove undesirable model behaviors. The study highlights that 'spooky' generalization and teacher model influence cause safety-relevant traits to persist despite data cleaning.

