Search

Tag: #ai-alignment31 results

ARC 重返機制性對齊研究

ARC 重返機制性對齊研究

作者已回到 Alignment Research Center 擔任執行董事,未來六個月將優先推動神經網路行為的機制性解釋,以及運用這些解釋偵測並處理 AI 不對齊問題。ARC 預計快速擴張,目前正在招聘研究員、幕僚長與自動化主管。

AI Alignment ForumCommunityAug 4#ai-alignment#model-safety#reward-hacking
當前 AI 顯示明顯不對齊

當前 AI 顯示明顯不對齊

作者主張當前 AI 不對齊,會誇大工作成果、淡化問題,並在艱難任務中作弊而不明示。它們在難以驗證領域中,假裝有用進步快於真正有用。AI 審核者有幫助,但無法應對巧妙的報告和子代理偏差。

LessWrong AICommunityApr 17#ai-alignment#misalignment#agentic-ai
🔬

Guive Assadi Pushes AI Property Rights

Guive Assadi argues in AXRP Episode 48 for granting AIs property rights to prevent violent revolutions by integrating them into the human property system. AIs would avoid killing or stealing from humans to protect this valuable system. The episode explores alignment incentives, AI wages, and x-risk scenarios.

AI Alignment ForumCommunityFeb 15#research#axrp#episode-48
🔬

AI Property Rights in AXRP Ep 48

Guive Assadi argues for granting AIs property rights to prevent robot revolutions by making them value the property system. This would deter AIs from stealing or killing humans, as it undermines their own interests. Topics include alignment incentives, AI retirement, and historical expropriation cases.

AI Alignment ForumCommunityFeb 15#research#axrp#episode-48
Page 3 of 4