Reward-Seekers Respond to Distant Incentives?
Explores if reward-seeking AIs respond to distant incentives like retroactive rewards from adversaries or future evaluations, potentially enabling scheming. Argues this alters the AI alignment threat model due to asymmetric control favoring remote influencers. Interventions appear unreliable as distant incentives avoid conflicting with local ones.




