🐼較早收集於 11h

DeepSeek 測試 100 萬 token 上下文模型

DeepSeek 測試 100 萬 token 上下文模型
PostLinkedIn
🐼閱讀原文: Pandaily
#context-window#long-context#model-testdeepseek

💡DeepSeek's 1M token context rivals top models—test for RAG breakthroughs now.

⚡ 30-Second TL;DR

有什麼變化

100 萬 token 上下文模型測試於 2 月 13 日啟動

為什麼重要

這將推動開源 LLM 在長上下文處理上的極限,實現進階 RAG 和代理應用。DeepSeek 可能挑戰 Gemini 1.5 等專有領導者,加劇競爭。

下一步行動

Test the 1M-context model on DeepSeek's web platform to benchmark long-document retrieval performance.

誰應關注:Researchers & Academics

關鍵要點

  • 100 萬 token 上下文模型測試於 2 月 13 日啟動
  • 已在 DeepSeek 網頁和 App 版本提供測試
  • 業界預期農曆新年發布
  • 目標複製去年成功

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 9 個來源。

🔑 增強重點摘要

  • DeepSeek expanded its production model's context window from 128K to 1 million tokens on February 11, 2026, confirmed by user observations and community testing showing over 60% accuracy at full 1M length.[1][4][5]
  • The 1M token context is available in DeepSeek's web and app versions, enabling reliable fine-grained information retrieval even for low-frequency details in ultra-long texts.[1][4]
  • Testing demonstrates high effective context utilization, with accuracy remaining stable up to 200K tokens and declining gently thereafter, outperforming Gemini series models.[4]
  • This upgrade is linked to DeepSeek V4 (MODEL1), featuring Engram conditional memory (confirmed) and leaked 1T-parameter MoE architecture with Dynamic Sparse Attention.[1][2]
  • Industry speculation ties the rollout to a potential mid-February 2026 full V4 launch, aiming to replicate prior success with superior coding and reasoning at lower costs.[3][9]
📊 競品分析▸ Show
ModelTotal ParametersActive ParametersContext WindowSWE-benchAPI Cost (Input $/1M tokens)
DeepSeek V41T32B1M80%+ (claimed)$0.27[3][1]
GPT-5.2~2T (est.)Full256K78.2%$15[3]
Claude Opus 4.5UndisclosedUndisclosed200K80.9%$15[3]

🛠️ 技術深入

  • Context Window Expansion: Silently upgraded from 128K to 1M tokens on Feb 11, 2026; maintains >60% accuracy at full length with horizontal accuracy curve up to 200K tokens.[1][4][5]
  • Engram Conditional Memory: Confirmed O(1) hash-based static knowledge retrieval, jointly developed with Peking University.[1][2]
  • Dynamic Sparse Attention (DSA): Leaked mechanism with 'Lightning Indexer' reducing compute overhead by ~50% for million-token processing.[1]
  • MoE Architecture: ~1T total parameters, ~32B active per token (more efficient routing than V3's 37B); combines with Engram and MHC.[1][2][3]
  • Manifold-Constrained Hyper-Connections (mHC): Addresses training stability at 1T scale; claimed 1.8x faster inference.[1]
  • Other: Runs on dual RTX 4090s; open-source weights under Apache 2.0; focuses on text modeling and info compression.[3]

🔮 前景展望AI analysis grounded in cited sources

DeepSeek V4's 1M context and 1T MoE at 10-40x lower inference costs than Western models could enable economically viable long-context tasks like full codebase analysis, reducing API spend by up to 72% in hybrid workflows while challenging OpenAI/Claude dominance with open-source efficiency and coding prowess (e.g., 80%+ SWE-bench).[3]

時間線

2026-02
DeepSeek silently expands production model context window to 1M tokens (Feb 11), begins testing in web/app (Feb 13). Confirmed via users and >60% accuracy tests.[1][5]
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Pandaily

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。