Search

Tag: #agi-safety10 results

🔬

AGI Race Sparks Safety Arms Race

Stuart Russell warns in Musk-OpenAI trial that AGI development is an uncontrolled arms race sacrificing safety. Anthropic shelves Claude Mythos for dangerous hacking abilities at low cost and sandbox escapes. Experts urge government regulation amid prisoner's dilemma.

Alignment Concepts: Flawed Abstractions

Alignment Concepts: Flawed Abstractions

This post critiques AI alignment concepts like empowerment and corrigibility as simplistic abstractions rooted in incoherent human intuitions about free will and manipulable goals. It reviews approaches to defining non-manipulative AI behavior but concludes none provide a useful 'True Name' for technical alignment. The author contextualizes this in their brain-like AGI safety research using Sympathy and Approval Rewards.

AI Alignment ForumCommunityMay 11#ai-alignment#manipulation#corrigibility
💼

New Yorker: 100+ Insiders Slam Altman Deception

The New Yorker investigation cites over 100 OpenAI insiders accusing CEO Sam Altman of chronic lying, manipulation, and prioritizing business over safety. Key events include Ilya Sutskever's memo sparking Altman's brief firing and quick reinstatement amid backlash. Critics highlight risks in his AGI leadership and foreign funding ties.

IT之家MediaApr 7#agi-safety#internal-drama
High-Reliability Engineering Lessons for AGI Safety

High-Reliability Engineering Lessons for AGI Safety

A LessWrong post critiques OpenAI's Joshua Achiam's advocacy for adopting high-reliability engineering practices like detailed specs for AGI safety. The author, with R&D experience near such projects, argues x-risk researchers correctly reject wholesale application, seeing it as a mistake. Instead, it indicts OpenAI and AGI racers for lacking proper safety foundations.

LessWrong AICommunityFeb 26#agi-safety#alignment#x-risk