預期無情反社會ASI為預設
文章辯論為何基於演員-評論者RL代理的腦狀AGI預設會是無情反社會,不同於LLMs或人類。作者反對現有心智的先驗,因為AI架構差異極大。強調專注特定RL-AGI威脅模型,而非一般智能類比。
Tag: #alignment84 results
文章辯論為何基於演員-評論者RL代理的腦狀AGI預設會是無情反社會,不同於LLMs或人類。作者反對現有心智的先驗,因為AI架構差異極大。強調專注特定RL-AGI威脅模型,而非一般智能類比。
LLMs lack human-like metacognitive skills for error-catching and cognition management. Enhancing these could cut slop, sycophancy, and aid alignment research. Benefits for alignment may outweigh capability risks.
Deferring to capable AIs is inevitable for risk management as control fails. Focus on minimal capability threshold, wise epistemics, and alignment on messy tasks. Prosaic methods for rushed scenarios using supervised AI labor.

OpenAI has disbanded its mission alignment team. The leader transitions to chief futurist role. Remaining members reassigned across the company.