Search

Tag: #interpretability59 results

CRL Steers SAE Features Token-by-Token

CRL Steers SAE Features Token-by-Token

CRL uses reinforcement learning to select sparse autoencoder (SAE) features for steering language models at each token, revealing which features impact outputs. It includes adaptive masking for diverse features and enables analysis like branch point tracking and layer-wise comparisons. Tested on Gemma-2 2B, it improves benchmarks while providing interpretable logs.

ArXiv AIResearchFeb 12#research#crl#gemma-2
Adapters Unlock Reliable Self-Interpretation

Adapters Unlock Reliable Self-Interpretation

Lightweight adapters trained on interpretability artifacts enable reliable self-interpretation in frozen LMs. A simple scalar affine adapter outperforms baselines in feature labeling, topic identification, and implicit reasoning decoding. Gains scale with model size, driven mostly by learned bias.

ArXiv AIResearchFeb 12#research#self-interpretation#v1
Page 6 of 6