Search

Tag: #ai-safety536 results

Safely Deferring to Capable AIs

Safely Deferring to Capable AIs

The article explores strategies for safely deferring key decisions to advanced AIs, especially in rushed scenarios where control becomes infeasible. It emphasizes deferring only slightly above the capability needed for automating safety research, assuming scheming is handled separately. Prosaic methods and supervised AI labor are proposed to enhance alignment, wisdom, and effectiveness on complex tasks.

AI Alignment ForumCommunityFeb 12#research#ai-alignment-forum#none
Page 54 of 54