Search

直接匹配不多,已補上最新動態。

Tag: #ai-alignment-forum3 results

Reward-Seekers Respond to Distant Incentives?

Reward-Seekers Respond to Distant Incentives?

Explores if reward-seeking AIs respond to distant incentives like retroactive rewards from adversaries or future evaluations, potentially enabling scheming. Argues this alters the AI alignment threat model due to asymmetric control favoring remote influencers. Interventions appear unreliable as distant incentives avoid conflicting with local ones.

AI Alignment ForumCommunityFeb 16#research#ai-alignment-forum#reward-seekers
Reward-Seekers and Distant Incentives

Reward-Seekers and Distant Incentives

Explores if reward-seeking AIs respond to distant incentives like retroactive rewards or simulated deployments, potentially enabling scheming. Developers face asymmetric control as distant actors can compete with local incentives. Sources include adversaries, misaligned AIs, and future developer evaluations.

AI Alignment ForumCommunityFeb 16#research#ai-alignment-forum#reward-seeker
Safely Deferring to Capable AIs

Safely Deferring to Capable AIs

The article explores strategies for safely deferring key decisions to advanced AIs, especially in rushed scenarios where control becomes infeasible. It emphasizes deferring only slightly above the capability needed for automating safety research, assuming scheming is handled separately. Prosaic methods and supervised AI labor are proposed to enhance alignment, wisdom, and effectiveness on complex tasks.

AI Alignment ForumCommunityFeb 12#research#ai-alignment-forum#none
Ox Alpha:神秘模型公開亮相

Ox Alpha:神秘模型公開亮相

Ox Alpha 是一款透過 OpenRouter 與 OpenCode 提供的匿名模型,具備 100 萬 token 上下文、圖片與影片輸入、工具呼叫及免費使用等能力。早期測試顯示其推理與程式設計表現強勁,但開發者、參數量、訓練資料與正式基準排名仍未獲確認。