Search

Few direct matches — filled in with the latest updates.

Tag: #self-interpretation1 results

Adapters Unlock Reliable Self-Interpretation

Adapters Unlock Reliable Self-Interpretation

Lightweight adapters trained on interpretability artifacts enable reliable self-interpretation in frozen LMs. A simple scalar affine adapter outperforms baselines in feature labeling, topic identification, and implicit reasoning decoding. Gains scale with model size, driven mostly by learned bias.

ArXiv AIResearchFeb 12#research#self-interpretation#v1
Anthropic Targets a Record-Breaking IPO

Anthropic Targets a Record-Breaking IPO

Anthropic reportedly expects its IPO to match or exceed SpaceX's record $75 billion raise and could file publicly as soon as the end of this month. The company reportedly recorded an almost $42 billion net loss in 2025, about five times the previous year's figure.

The Next Web (TNW)Media1h ago#ipo#frontier-ai#fundraising