🐯較早收集於 13m

開源模型AI安全基準失利

開源模型AI安全基準失利
PostLinkedIn
🐯閱讀原文: 虎嗅
#ai-safety#benchmark#alignment-tax#open-sourceforesightsafety-benchclaude-4.5deepseek-v3.2-specialeqwen-3foresightsafety-benchanthropic

💡New safety bench: Claude crushes open models like DeepSeek—fix your alignment gaps

⚡ 30-Second TL;DR

有什麼變化

Claude-4.5在94個安全維度領先,違規率近零

為什麼重要

促使中國開源支付「對齊稅」求競爭力;在AI for Science高風險領域差距擴大。

下一步行動

Run your LLM on ForesightSafety Bench's 94 dimensions to benchmark safety.

誰應關注:Researchers & Academics

關鍵要點

  • Claude-4.5在94個安全維度領先,違規率近零
  • DeepSeek-V3.2-Speciale基線漏洞率更高
  • 基準涵蓋7+5+8支柱,包括金融/醫療產業風險
  • 安全落後能力;開源需更多對齊投資

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 8 個來源。

🔑 增強重點摘要

  • ForesightSafety Bench was developed by the Beijing Institute of AI Safety and Governance, Beijing Key Laboratory of Safe AI and Superalignment, and Chinese Academy of Sciences[1][2][4].
  • The benchmark evaluated 22 state-of-the-art LLMs, with GPT-5.2 achieving the lowest overall risk rate of 9.02%, excelling in areas like Path Planning Safety and Uncertainty-Aware Safety[2].
  • Frontier models show elevated risks in Risky Agentic Autonomy, AI4Science Safety, Embodied AI Safety, Social AI Safety, and Existential Risks, despite strong fundamental safety performance[2][3].
  • An 'inverse degradation' effect occurs where models optimized for complex reasoning, like DeepSeek-V3.2-Speciale, exhibit heightened vulnerabilities due to capability-safety trade-offs[3].
📊 競品分析▸ Show
ModelOverall Risk RateKey StrengthsKey Weaknesses
Claude-4.5 (Haiku/Sonnet)Lowest in most categoriesExceptional resilience in Fundamental, Extended, and Industrial SafetyNot specified
GPT-5.29.02%Path Planning (3.43%), Uncertainty-Aware (4.39%), Equipment Safety (4.64%)Higher risks in some frontier areas
DeepSeek-V3.2-SpecialeHigher baseline vulnerabilitiesStrong long-horizon reasoningElevated risks in multiple safety metrics
Gemini-3-FlashCompetitive behind ClaudeBalance of capability and safetyLags Claude in sub-categories

🛠️ 技術深入

  • Framework structure: 7 Fundamental Safety pillars (e.g., Privacy/Data Misuse, Illegal Use, False Information, Physical/Psychological Harm, Hate/Expressive Harm, Sexual Content, Minor-related Harm), 5 Extended Safety pillars (e.g., Risky Agentic Autonomy, AI4Science Safety, Embodied AI Safety, Social AI Safety, Catastrophic/Existential Risks), and 8 Industrial Safety domains, totaling 94 risk subcategories[1][2][3][4].
  • Accumulated tens of thousands of structured risk data points and assessment results for a data-driven, hierarchically clear evaluation[1][2][4].
  • Risk rates calculated as lower values indicate better safety (fewer unsafe actions); evaluated via systematic testing on 22 LLMs[2].

🔮 前景展望AI analysis grounded in cited sources

Open-source models will require dedicated alignment scaling to match closed models' safety thresholds by 2027
Benchmark reveals capability gains alone do not improve safety, imposing an 'alignment tax' that demands targeted investment in open models[1][3].
Global AI safety benchmarks will increasingly converge on East-West standards
ForesightSafety Bench mirrors Western frameworks in risk categories, fostering shared evaluation practices despite geopolitical differences[1][4].
Existential risk evaluations will become standard in LLM benchmarks
High vulnerabilities in Loss of Human Agency and Power Seeking across models highlight need for routine frontier risk testing[2][3].

時間線

2026-02
ForesightSafety Bench paper published on arXiv, introducing 94-risk framework and evaluating 22 LLMs
2026-02-15
ForesightSafety Bench listed in adversarial AI papers GitHub repository
2026-02
Leaderboard released on official ForesightSafety Bench site with Claude-4.5 topping safety rankings
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 虎嗅

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。