📊較早收集於 4h

AI影像先驅新創Inception推出加速聊天技術

AI影像先驅新創Inception推出加速聊天技術
PostLinkedIn
📊閱讀原文: Bloomberg Technology

💡Image AI pioneer's text tech promises faster LLM chats

⚡ 30-Second TL;DR

有什麼變化

Stefano Ermon的Inception推出聊天加速技術

為什麼重要

可能改善基於LLM的聊天延遲,有助即時AI應用。

下一步行動

Test Inception's demo for text speedups in your chatbot prototypes.

誰應關注:Developers & AI Engineers

關鍵要點

  • Stefano Ermon的Inception推出聊天加速技術
  • 先前開創AI影像/影片生成
  • 現瞄準更快的AI文字處理

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 6 個來源。

🔑 增強重點摘要

  • Inception Labs applies diffusion models to language modeling, generating text in parallel rather than sequentially like autoregressive transformers, achieving over 1000 tokens/sec on NVIDIA H100 GPUs[1][2][5].
  • Mercury family includes Mercury Coder for code generation and a general-purpose model, both with 128,000-token context windows, priced at 25 cents per million input tokens and $1 per million output tokens[2].
  • Company raised $50M in seed funding led by Menlo Ventures, with investors including Andrew Ng and Andrej Karpathy, to scale diffusion LLMs[2][4][6].

🛠️ 技術深入

  • Uses diffusion process to generate entire blocks of text at once via denoising steps, retaining transformer neural architecture but with different training objective and parallel inference[1][3].
  • Early prototype matched GPT-2 quality at 10x faster speed; scaled up with larger models and better data for commercial Mercury[1].
  • Features animation in chat interface visualizing text sharpening from noise to detail during generation[2].
  • Reduces GPU footprint, enabling larger models at same latency/cost or more users on existing infrastructure[2].

🔮 前景展望AI analysis grounded in cited sources

Diffusion LLMs will enable larger models in latency-sensitive apps without sacrificing speed
dLLMs allow drop-in replacement of autoregressive models, letting partners use more capable models while meeting original cost and latency requirements, as reported by early adopters in customer support and automation[5].
dLLMs reduce inference costs for long reasoning traces
Parallel generation avoids sequential token-by-token processing, countering ballooning costs from test-time computation in frontier autoregressive LLMs[5].

時間線

2019-01
Stefano Ermon invents diffusion models for text at Stanford lab
2025-02
Inception launches first commercial dLLM, Mercury
2025-11
Inception raises $50M seed round led by Menlo Ventures
2025-12
Public interview reveals Mercury Coder and diffusion LLM details
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Bloomberg Technology

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。