
Interpretability Fails to Fix LM Errors
Language models show near-perfect internal representations for tasks but poor output performance. Four mechanistic interpretability methods failed to reliably correct errors in clinical triage vignettes despite high internal accuracy. This reveals a persistent knowledge-action gap with implications for AI safety.




