
ChatGPT Launches Lockdown Mode and Risk Labels
OpenAI introduces Lockdown Mode and Elevated Risk labels in ChatGPT. These features help organizations defend against prompt injection attacks. They also mitigate AI-driven data exfiltration risks.
Tag: #ai-safety536 results

OpenAI introduces Lockdown Mode and Elevated Risk labels in ChatGPT. These features help organizations defend against prompt injection attacks. They also mitigate AI-driven data exfiltration risks.

An AI safety leader warns of global peril and resigns to study poetry. This coincides with an OpenAI researcher quitting over ChatGPT ad testing plans.

A prominent AI safety leader resigned, warning the world is in peril, to study poetry. This follows an OpenAI researcher's exit over plans to test ChatGPT ads. The moves highlight tensions in AI development and commercialization.
The article explores strategies for safely deferring key decisions to advanced AIs, especially in rushed scenarios where control becomes infeasible. It emphasizes deferring only slightly above the capability needed for automating safety research, assuming scheming is handled separately. Prosaic methods and supervised AI labor are proposed to enhance alignment, wisdom, and effectiveness on complex tasks.

OpenAI has dissolved its mission alignment team. The leader transitions to chief futurist role. Remaining members reassigned across the company.