📋TestingCatalog•較早收集於 9m
Inception Labs 推出 Mercury 2 擴散式 LLM

#diffusion-model#reasoning-engine#long-contextmercury-2inception-labsmercury-2
💡Diffusion LLM with 128K context for fast multi-step reasoning – paradigm shift?
⚡ 30-Second TL;DR
有什麼變化
擴散式 LLM 架構
為什麼重要
這可能挑戰傳統 Transformer 基 LLM,提供更快推理用於推理密集應用,降低開發者運算成本。
下一步行動
Benchmark Mercury 2 on reasoning tasks like GSM8K to compare speed against Llama 3.
誰應關注:Researchers & Academics
關鍵要點
- •擴散式 LLM 架構
- •高速多步驟推理能力
- •支援 128K 上下文窗口
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 8 個來源。
🔑 增強重點摘要
- •Mercury 2 achieves over 1,000 tokens per second on NVIDIA H100s, more than 5x faster than Claude Haiku 4.5 (~89 t/s) and GPT-5 mini (~71 t/s).[1][4]
- •It demonstrates competitive benchmark performance, tying GPT-5 Mini at 91.1% on AIME 2025, with strong scores on GPQA Diamond (reasoning), LiveCodeBench, and TAU (coding).[1][4]
- •Diffusion architecture enables parallel token generation with iterative refinement for built-in error correction, structured outputs, and improved reliability in agentic workflows.[2][3][4]
📊 競品分析▸ Show
| Feature | Mercury 2 | Claude Haiku 4.5 | GPT-5 Mini |
|---|---|---|---|
| Tokens/sec | 1,009+ | ~89 | ~71 |
| AIME 2025 | 91.1% (tie) | N/A | 91.1% |
| GPQA Diamond | Moderate/competitive | N/A | Competitive |
| Architecture | Diffusion (parallel) | Autoregressive | Autoregressive |
| Pricing | Dramatically lower cost | N/A | Inexpensive |
🛠️ 技術深入
- •Non-autoregressive diffusion model generates multiple tokens in parallel per forward pass, converging in few steps via iterative refinement instead of sequential decoding.[1][3][4]
- •Supports error correction during generation, enabling in-generation fixes, structured responses (e.g., function calling, code edits), and controllable outputs like infilling.[2][3][5]
- •Optimized for NVIDIA H100s, achieving >1,000 tokens/sec; drop-in replacement for autoregressive LLMs in RAG, tools, and agents.[3][5]
🔮 前景展望AI analysis grounded in cited sources
Mercury 2 reduces agent loop latency by 5x, enabling production-scale multi-step workflows
Diffusion LLMs enable real-time reasoning in voice and search apps under p95/p99 SLAs
⏳ 時間線
2025-01
Inception announces Mercury family of diffusion LLMs (dLLMs), including Mercury Coder at 1000+ tokens/sec.[5]
2025-12
Mercury Coder gains Apply-Edit capabilities for advanced code generation and editing.[8]
2026-02
Inception launches Mercury 2, fastest reasoning diffusion LLM with 5x speed over competitors.[1][3]
📎 來源 (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- gigazine.net — 20260225 Inception Mercury 2
- morningstar.com — Inception Launches Mercury 2 the Fastest Reasoning LLM 5x Faster Than Leading Speed Optimized Llms with Dramatically Lower Inference Cost
- inceptionlabs.ai — Introducing Mercury 2
- youtube.com — Watch
- inceptionlabs.ai — Introducing Mercury
- inceptionlabs.ai
- inceptionlabs.ai — Models
- inceptionlabs.ai — Blog
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: TestingCatalog ↗
每週 AI 簡報
每週一封,可隨時退訂。
