Search

Few direct matches — filled in with the latest updates.

Tag: #language-models4 results

Interpretability Fails to Fix LM Errors

Interpretability Fails to Fix LM Errors

Language models show near-perfect internal representations for tasks but poor output performance. Four mechanistic interpretability methods failed to reliably correct errors in clinical triage vignettes despite high internal accuracy. This reveals a persistent knowledge-action gap with implications for AI safety.

Tokens Enable Emergent Resource Rationality

Tokens Enable Emergent Resource Rationality

Inference-time scaling in language models leads to adaptive resource rationality without explicit cost rewards. Models shift from brute-force to analytic strategies as task complexity rises. LRMs show robustness on challenging functions like XOR/XNOR unlike IT models.

ArXiv AIResearchFeb 12#research#language-models#v1
Anthropic Targets a Record-Breaking IPO

Anthropic Targets a Record-Breaking IPO

Anthropic reportedly expects its IPO to match or exceed SpaceX's record $75 billion raise and could file publicly as soon as the end of this month. The company reportedly recorded an almost $42 billion net loss in 2025, about five times the previous year's figure.

The Next Web (TNW)Media51m ago#ipo#frontier-ai#fundraising