Anthropic 發布 Sonnet 4.6

💡Anthropic's mid-size LLM update keeps pace—test for better perf/cost balance (62 chars)
⚡ 30-Second TL;DR
有什麼變化
Anthropic 推出 Sonnet 4.6 模型
為什麼重要
此發布強化 Anthropic 在中型 LLM 的地位,為使用者提供潛在改進效能,而無需等待更久。AI 從業者可整合它,用於相較大型模型更具成本效益的推理。
下一步行動
Test Sonnet 4.6 via Anthropic API on your mid-size model benchmarks today.
關鍵要點
- •Anthropic 推出 Sonnet 4.6 模型
- •針對中型模型領域
- •維持四個月更新週期
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 5 個來源。
🔑 增強重點摘要
- •Claude Sonnet 4.6 achieves 72.5 on OSWorld-Verified benchmark, up from 28.0 for Sonnet 3.7, demonstrating major improvements in computer use automation capabilities
- •Sonnet 4.6 delivers performance previously requiring Opus-class models on real-world office tasks like spreadsheet navigation and multi-step web forms, narrowing the capability gap between mid-tier and premium models
- •Model features 1M token context window in beta and 200K standard context window, with 64K max output tokens and support for extended thinking and adaptive thinking
- •Anthropic upgraded free-tier Claude users to Sonnet 4.6 by default with file creation, connectors, skills, and context compaction included, expanding accessibility
- •Enhanced safety measures show Sonnet 4.6 demonstrates major improvement in prompt injection resistance compared to Sonnet 4.5, performing similarly to Opus 4.6
📊 競品分析▸ Show
| Aspect | Claude Sonnet 4.6 | Claude Opus 4.6 | Notes |
|---|---|---|---|
| Context Window | 1M (beta) / 200K standard | 1M (beta) / 200K standard | Both support extended context |
| Max Output Tokens | 64K | 128K | Opus maintains higher output capacity |
| Primary Use Case | Speed-intelligence balance | Maximum capability, agentic tasks | Sonnet targets broader user base |
| Thinking Modes | Extended, Adaptive | Extended, Adaptive | Both support reasoning enhancements |
| Availability | All plans including free tier | Pro/Max/Team/API | Sonnet more accessible |
| Computer Use Benchmark | 72.5 (OSWorld-Verified) | Not separately specified | Sonnet shows significant improvement trajectory |
🛠️ 技術深入
• Context Window: Supports 200K tokens standard with 1M token context window available in beta; context compaction feature automatically summarizes older context during long conversations • Output Capacity: 64K maximum output tokens for structured responses • Thinking Capabilities: Supports both extended thinking (deliberative reasoning) and adaptive thinking (contextual reasoning adjustment) • Effort Parameter: Introduces effort levels (low, medium, high, max) allowing developers to balance speed, cost, and performance • Web Tools: Web search and fetch tools now automatically write and execute code to filter and process results, improving token efficiency • Tool Availability: Code execution, memory, programmatic tool calling, tool search, and tool use examples now generally available on API • Safety Architecture: Demonstrates improved resistance to prompt injections with behavioral audits showing emotional stability metrics • Coding Improvements: Enhanced consistency, instruction following, and code review capabilities; developers with early access prefer it over Opus 4.5 from November 2025
🔮 前景展望AI analysis grounded in cited sources
Sonnet 4.6's performance parity with Opus-class models on economically valuable office tasks suggests a flattening of capability tiers, potentially disrupting premium pricing models. The 72.5 OSWorld benchmark score represents a 2.6x improvement over Sonnet 3.7, indicating accelerating progress in agentic computer use—a critical capability for autonomous task automation. Expanded free-tier access with advanced features (file creation, connectors, compaction) may drive broader adoption and developer ecosystem growth. The emphasis on safety improvements and prompt injection resistance addresses enterprise deployment concerns. Anthropic's consistent four-month release cadence and rapid capability improvements position the company to maintain competitive pressure against OpenAI and other AI providers in the mid-market segment, where cost-performance tradeoffs are critical.
⏳ 時間線
📎 來源 (5)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: TechCrunch AI ↗
每週 AI 簡報
每週一封,可隨時退訂。



