🐯較早收集於 5m

Anthropic 揭露中國 Claude 蒸餾攻擊

Anthropic 揭露中國 Claude 蒸餾攻擊
PostLinkedIn
🐯閱讀原文: 虎嗅

💡Chinese labs distilled Claude at scale—learn evasion tactics & defenses for your APIs

⚡ 30-Second TL;DR

有什麼變化

DeepSeek:超過15萬聊天,針對推理、RL 獎勵、審查改述

為什麼重要

暴露 API 對蒸餾漏洞,可能促使前沿模型更嚴格速率限制與地理封鎖。加劇蒸餾倫理與行業慣例辯論。

下一步行動

Audit your API prompts for distillation patterns using metadata like IP clustering and rubric repeats.

誰應關注:Researchers & Academics

關鍵要點

  • DeepSeek:超過15萬聊天,針對推理、RL 獎勵、審查改述
  • Moonshot AI (Kimi):超過340萬,針對代理、編程、電腦視覺
  • MiniMax:超過1300萬,代理編程、工具;適應新 Claude 版本

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 5 個來源。

🔑 增強重點摘要

  • Anthropic attributed the campaigns to specific labs using IP address correlation, request metadata, infrastructure indicators, and corroboration from industry partners.[1][4]
  • Anthropic has deployed classifiers, behavioral fingerprinting systems, strengthened verification for educational and startup accounts, and enhanced safeguards to counter distillation attacks.[3][4]
  • OpenAI has previously accused DeepSeek of similar distillation techniques on its models.[2]
  • Google Threat Intelligence Group recently disrupted distillation attacks targeting Gemini's reasoning capabilities via over 100,000 prompts.[3]
  • The attacks occurred amid US debates on AI chip exports, with the Trump administration allowing Nvidia H200 chips to China last month.[1]

🛠️ 技術深入

  • Distillation involves sending bulk structured prompts at industrial scale to extract capabilities like agentic reasoning traces, which attackers later reconstruct.[4]
  • Moonshot AI's later phase specifically attempted to extract and reconstruct Claude’s internal reasoning traces through targeted prompts.[4]
  • Detection relies on classifiers identifying anomalous API traffic patterns distinct from normal usage, including volume, structure, and prompt focus.[4]
  • Attribution methods include IP correlation across proxy services (hydra clusters), request metadata matching staff profiles, and shared infrastructure indicators.[1][4]

🔮 前景展望AI analysis grounded in cited sources

Distilled models will proliferate without US safety safeguards, enabling authoritarian cyber operations.
Anthropic states illicitly distilled models strip protections against bioweapons, cyber activities, disinformation, and surveillance when fed into foreign military systems.[1][2][4]
AI industry will adopt coordinated defenses including cloud providers and policymakers.
Anthropic calls for collective response beyond individual investments, as distillation circumvents export controls preserving US AI lead.[1][4]
Distillation attacks will target all frontier models like Gemini and OpenAI systems.
Similar campaigns hit Google Gemini and OpenAI, with patterns adapting across providers using fraudulent accounts and proxies.[2][3]

時間線

2026-01
Trump administration allows Nvidia H200 AI chip exports to China amid debates.
2026-02
Anthropic publishes blog exposing DeepSeek, Moonshot AI, and MiniMax distillation campaigns on Claude.

📰 事件追蹤

📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 虎嗅

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。