📡TechRadar AI•較早收集於 55m
Anthropic 取消安全承諾,重寫 AI 護欄

💡Anthropic's safety pivot speeds model releases but drops key safeguards—critical for AI devs.
⚡ 30-Second TL;DR
有什麼變化
重寫旗艦安全政策
為什麼重要
此政策轉變可能加速 Anthropic 的模型發布,從而影響產業安全標準。從業人員應評估依賴 Anthropic 模型的風險,因為剛性保障減少。
下一步行動
Review Anthropic's updated Responsible Scaling Policy on their site for deployment criteria changes.
誰應關注:Researchers & Academics
關鍵要點
- •重寫旗艦安全政策
- •取消「未安全不發布」承諾
- •轉向靈活的透明度框架
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 6 個來源。
🔑 增強重點摘要
- •Anthropic cited an 'anti-regulatory political climate' and lack of federal AI policy progress as key drivers for the policy rewrite, acknowledging that state-level efforts (California SB 53, New York RAISE Act) and international frameworks (EU AI Act) have created new compliance requirements that their RSP now addresses through public documentation including a Frontier Compliance Framework[2].
- •The new RSP introduces a 'Frontier Safety Roadmap' requirement—a public-facing document detailing concrete risk mitigation plans across Security, Alignment, Safeguards, and Policy—designed to maintain the incentive structure of the original policy by creating forcing functions for safety development[1][3].
- •Anthropic's ASL-3 safeguards (targeting risks from chemical and biological weapons) have been operationalized since May 2025 and proved feasible in practice, demonstrating that the company successfully developed sophisticated input/output classifiers to block harmful content, which informed the decision to maintain rather than eliminate safety standards[3].
- •The policy change reflects a strategic shift from unilateral constraint to industry-wide transparency standards, as Anthropic acknowledged its original RSP failed to persuade competitors to adopt similar 'pause scaling' commitments, creating competitive disadvantage against OpenAI, Microsoft, and other frontier labs[5].
🛠️ 技術深入
- •ASL-3 safeguards operationalized May 2025: input/output classifiers designed to block chemical/biological weapons content; access controls for trusted users with exemptions; red-teaming, bug bounties, and threat intelligence for jailbreak assessment; security controls with evolving methods but maintained rigor[3][4].
- •Constitutional Classifiers: core technical element of ASL-3 protections, with Anthropic committing to maintain or improve robustness at least equivalent to initial implementation[4].
- •Frontier Safety Roadmap framework: structured across four technical domains (Security, Alignment, Safeguards, Policy) with specific, measurable goals; example goal targets rare or jailbreak-requiring Constitutional violations on production Claude releases[4].
🔮 前景展望AI analysis grounded in cited sources
Anthropic will face pressure to expand ASL-3 protections beyond chemical/biological weapons vectors if it determines AI capabilities enable catastrophic threats in additional domains.
The RSP explicitly commits to applying protections 'at least as strong as current ASL-3 protections' to expanded use cases if new threat pathways emerge[4].
Federal AI regulation may become necessary to prevent competitive safety degradation across the industry, as Anthropic's unilateral policy shift demonstrates market failure in voluntary safety coordination.
METR policy director Chris Painter characterized the change as evidence that 'society is not prepared for potential catastrophic risks' and that risk assessment methods are not keeping pace with capabilities[1].
Anthropic's Frontier Safety Roadmaps will become a de facto industry standard for AI safety transparency, as governments (California, New York, EU) increasingly mandate catastrophic risk frameworks.
The RSP update explicitly notes that regulatory requirements for frontier AI developers to publish risk management frameworks are already in place, and Anthropic's public roadmap approach directly addresses these mandates[3].
⏳ 時間線
2023-01
Anthropic releases first version of Responsible Scaling Policy (RSP), establishing 'no release until safe' pledge as flagship safety commitment
2025-05
Anthropic activates ASL-3 safeguards for relevant models, operationalizing Constitutional Classifiers and access controls for chemical/biological weapons risk mitigation
2026-02
Anthropic releases RSP Version 3.0, dropping categorical pause-scaling pledge and introducing Frontier Safety Roadmap transparency framework
📎 來源 (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: TechRadar AI ↗
每週 AI 簡報
每週一封,可隨時退訂。
