📋較早收集於 9m

Inception Labs 推出 Mercury 2 擴散式 LLM

Inception Labs 推出 Mercury 2 擴散式 LLM
PostLinkedIn
📋閱讀原文: TestingCatalog
#diffusion-model#reasoning-engine#long-contextmercury-2inception-labsmercury-2

💡Diffusion LLM with 128K context for fast multi-step reasoning – paradigm shift?

⚡ 30-Second TL;DR

有什麼變化

擴散式 LLM 架構

為什麼重要

這可能挑戰傳統 Transformer 基 LLM,提供更快推理用於推理密集應用,降低開發者運算成本。

下一步行動

Benchmark Mercury 2 on reasoning tasks like GSM8K to compare speed against Llama 3.

誰應關注:Researchers & Academics

關鍵要點

  • 擴散式 LLM 架構
  • 高速多步驟推理能力
  • 支援 128K 上下文窗口

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 8 個來源。

🔑 增強重點摘要

  • Mercury 2 achieves over 1,000 tokens per second on NVIDIA H100s, more than 5x faster than Claude Haiku 4.5 (~89 t/s) and GPT-5 mini (~71 t/s).[1][4]
  • It demonstrates competitive benchmark performance, tying GPT-5 Mini at 91.1% on AIME 2025, with strong scores on GPQA Diamond (reasoning), LiveCodeBench, and TAU (coding).[1][4]
  • Diffusion architecture enables parallel token generation with iterative refinement for built-in error correction, structured outputs, and improved reliability in agentic workflows.[2][3][4]
📊 競品分析▸ Show
FeatureMercury 2Claude Haiku 4.5GPT-5 Mini
Tokens/sec1,009+~89~71
AIME 202591.1% (tie)N/A91.1%
GPQA DiamondModerate/competitiveN/ACompetitive
ArchitectureDiffusion (parallel)AutoregressiveAutoregressive
PricingDramatically lower costN/AInexpensive

🛠️ 技術深入

  • Non-autoregressive diffusion model generates multiple tokens in parallel per forward pass, converging in few steps via iterative refinement instead of sequential decoding.[1][3][4]
  • Supports error correction during generation, enabling in-generation fixes, structured responses (e.g., function calling, code edits), and controllable outputs like infilling.[2][3][5]
  • Optimized for NVIDIA H100s, achieving >1,000 tokens/sec; drop-in replacement for autoregressive LLMs in RAG, tools, and agents.[3][5]

🔮 前景展望AI analysis grounded in cited sources

Mercury 2 reduces agent loop latency by 5x, enabling production-scale multi-step workflows
Parallel diffusion shrinks compounding delays in code agents, IT triage, and back-office automation, improving controllability and trust per Inception's deployment analysis.[2][3]
Diffusion LLMs enable real-time reasoning in voice and search apps under p95/p99 SLAs
Sub-second generation with reasoning quality supports natural UX in support agents, tutoring, and translation without retries.[2][3]
Iterative refinement boosts output reliability 20-30% in benchmarks vs. autoregressive models
Built-in error correction during parallel generation reduces hallucinations and fallbacks, as shown in GPQA and LiveCodeBench results.[1][4]

時間線

2025-01
Inception announces Mercury family of diffusion LLMs (dLLMs), including Mercury Coder at 1000+ tokens/sec.[5]
2025-12
Mercury Coder gains Apply-Edit capabilities for advanced code generation and editing.[8]
2026-02
Inception launches Mercury 2, fastest reasoning diffusion LLM with 5x speed over competitors.[1][3]
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: TestingCatalog

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。