
Research Directions for Training on Probes
The article proposes research directions for using probes to help AI systems generalize behavioral judgments from easy cases to harder domains. It explores weak supervision, teaching new skills, retraining probes as concepts evolve, runtime behavioral modulation, and alignment failures inspired by the brain.
AI Alignment Forum · 6d ago




















