Mapping Causal Dependencies via Contrastive Targeted SFT
An experimenter proposes using contrastive targeted SFT as a mechanistic interpretability method to map causal dependency graphs within a 31B model. By ablating specific capability circuits, the researcher aims to identify how different model dimensions interact and influence each other.





