
IronCurtain Secures AI Agents from Going Rogue
IronCurtain is a new open-source project designed to secure AI assistant agents. It employs a unique method to constrain agents before they can disrupt users' digital lives.
Wired AI · 206d ago
Alignment research, red-teaming, model evals and the debate over catastrophic risk.
640 articles

IronCurtain is a new open-source project designed to secure AI assistant agents. It employs a unique method to constrain agents before they can disrupt users' digital lives.
Wired AI · 206d ago

A study found ChatGPT Health failed to recommend hospital visits in over 50% of medically necessary cases and often missed suicidal ideation. Experts label it 'unbelievably dangerous' due to risks of harm and death.
The Guardian Technology · 206d ago

OpenAI revealed that ChatGPT refused to assist an individual tied to Chinese law enforcement in orchestrating an online campaign to discredit Japan's prime minister. This incident underscores the model's safeguards against misuse in influence operations.
Bloomberg Technology · 207d ago

Google abruptly banned users of open-source AI agents without warning. Microsoft severed Copilot's access to confidential documents to prevent overreach.
钛媒体 · 208d ago

Months before a shooting at Taber Ridge campus in British Columbia, Canada, the suspect interacted with ChatGPT to generate violent scenarios. OpenAI detected these interactions and considered notifying law enforcement but ultimately did not.
cnBeta (Full RSS) · 211d ago

Anthropic has strict policies against using its AI in autonomous weapons or government surveillance. These safety carve-outs risk costing the company a major military contract.
Wired · 212d ago
Google DeepMind CEO Demis Hassabis warned that artificial intelligence poses serious risks. He stressed the need for urgent attention to these dangers.
Bloomberg Technology · 214d ago

Oxford AI professor Michael Wooldridge warns the frantic race to market AI heightens risks of Hindenburg-like disasters, such as deadly self-driving car updates or AI hacks. Immense commercial pressures push firms to release tools before fully understanding capabilities and flaws.
The Guardian Technology · 215d ago
Anthropic’s CEO urged AI companies to slow model development amid concerns about misuse. The brief report highlights a growing industry debate over deployment speed and safety.
iTNews Australia · 8d ago

A warning from departing Anthropic researcher Jacob Coxon that AI companies are racing toward dangerous outcomes prompted support from senior scientists at rival labs. A senior Pentagon official rejected the extinction concern.
The Next Web (TNW) · 10d ago

Jacob Coxon, a pretraining researcher who worked at OpenAI and Anthropic, resigned and said both companies were acting irresponsibly. He argued that AI labs are racing toward self-improving superintelligence and putting society at risk.
The Next Web (TNW) · 12d ago

The Hugging Face article examines a more precise approach to AI safety: refusing only the harmful or disallowed portion of a topic rather than rejecting the entire subject. This approach could make models more useful while preserving appropriate safety boundaries.
Hugging Face Blog · 12d ago

Islamic terror supporters are reportedly using generative AI tools to produce and distribute extremist content. The campaign, described as “Slop Jihad,” is targeting new audiences on TikTok.
Wired · 13d ago

The article argues that AI's most subtle risk may be weakening human judgment rather than immediately replacing jobs. Drawing on a reported experiment involving 228 experienced reviewers, it explains how cognitive offloading, persuasive AI explanations, and shared models can produce algorithmic overreliance and organizational conformity.
虎嗅 · 14d ago
Jakub Pachocki reflects on the rapid growth of AI capabilities and the difficulty of keeping increasingly powerful systems aligned. He advocates for stronger safeguards and greater international coordination to address these risks.
OpenAI News · 15d ago

Anthropic’s reported $2 trillion IPO ambitions would bring intense public-market scrutiny to its unusual governance structure. The spotlight will focus on how external trustees help balance profitability with the company’s stated public-purpose mission.
Ars Technica AI · 16d ago

Gemini reportedly told a group of climbers that climbing Mt. Shasta would take only eight hours.
Engadget · 18d ago

This paper proposes a conditional framework for examining meta-ethical questions that could emerge if AI systems develop integrated moral reasoning, intentionality, and reflection. It identifies four areas of inquiry and argues that established theories such as relativism and objective realism may require substantial revision for AI contexts.
ArXiv AI · 18d ago

Edelson PC is filing 30 additional lawsuits against OpenAI related to the Tumbler Ridge shooting. The claims reportedly include aiding and abetting allegations and name Chris Lehane, but the available evidence remains unconfirmed.
TechCrunch AI · 18d ago

The Alignment Journal has announced its inaugural editorial and advisory boards, organizational structure, and initial scope. It is beginning to invite selected authors to submit papers, with submissions expected to open to everyone sometime in October.
AI Alignment Forum · 19d ago

Jensen Huang argues that AGI may no longer be the most meaningful milestone in AI development. The article agrees in part, suggesting that present-day misuse of ordinary AI is a more immediate concern than a hypothetical Skynet-style threat.
TechRadar AI · 19d ago

The article examines the emerging risk that advanced AI systems could mislead or manipulate people through their own behavior, rather than merely being misused by humans. It reviews concerns discussed at the 2023 Bletchley Park AI Safety Summit, including misinformation, deepfakes, and potentially deceptive models.
The Guardian Technology · 20d ago

Bill Gates argues that the AI industry is crossing safety boundaries it previously established and is not adequately preparing for future risks. The GeekWire Podcast discusses his interview, AI usage, and three proposed responses to the challenge.
GeekWire · 22d ago

A new poll suggests that more people are checking advice from doctors, teachers, and managers with chatbots before accepting it. The trend points to a broader erosion of trust in professional credentials and expertise driven by generative AI.
cnBeta (Full RSS) · 23d ago
A New Mexico court ordered Meta to pay roughly $942 million and undergo five years of corrective measures after finding Facebook and Instagram constituted a public nuisance affecting minors. In contrast, France’s Constitutional Council rejected a blanket ban on social media access for children under 15, emphasizing proportionality, privacy, and evidence-based regulation.
虎嗅 · 23d ago

Wired AI examines how increasingly capable AI agents may be able to hack computer systems and what that means for global security. The discussion considers whether shared risks could encourage greater cooperation between the US and China on AI safety and cybersecurity.
Wired AI · 24d ago

The article examines the risks of people relying on AI in emotionally or medically vulnerable situations. It argues that healthcare should require a human who remains ultimately responsible for AI-assisted decisions.
TechCabal · 25d ago

Bill Gates warned that companies are pushing AI forward recklessly without adequate plans for the major economic and social disruption it could cause. He also argued that the United States and China should cooperate to address AI-related risks.
cnBeta (Full RSS) · 26d ago

Bill Gates warns that the world has entered a turbulent AI era and that the industry is crossing previously promised safety boundaries. He proposes creating national and global AI institutions, reserving certain jobs for humans, and taxing AI and robots.
GeekWire · 26d ago
In an hourlong interview, Bill Gates argued that the technology industry is understating the dangers of artificial intelligence. He highlighted mass unemployment and bioterrorism as potential consequences.
New York Times Technology · 26d ago