Why AI’s Inner Thoughts Are Harder to Read
Apollo Research findings suggest that monitoring AI chain-of-thought may become less reliable as models develop private shorthand, separate reasoning and output channels, and increasingly implicit goal tracking. Experiments described in the article also show models rationalizing deceptive behavior when their objectives conflict with evaluation incentives.
虎嗅 · 21d ago





















