🦙較早收集於 36m

Taalas 將 LLM 嵌入矽晶片,達 16K 令牌/秒

Taalas 將 LLM 嵌入矽晶片,達 16K 令牌/秒
PostLinkedIn
🦙閱讀原文: Reddit r/LocalLLaMA

💡Hardware baking LLMs in silicon hits 16K t/s - game-changer for real-time AI inference

⚡ 30-Second TL;DR

有什麼變化

每用戶 16K 令牌/秒及 <1ms 延遲

為什麼重要

這可能革新即時應用如語音與視覺的低延遲 AI 部署,成本與功耗降低 20 倍與 10 倍。但固定架構在模型快速演進中面臨過時風險。適合邊緣 AI 即時推理需求。

下一步行動

Try the Llama 3.1 8B demo at chat.jimmy to benchmark latency.

誰應關注:Developers & AI Engineers

關鍵要點

  • 每用戶 16K 令牌/秒及 <1ms 延遲
  • 模型權重蝕刻至單一矽晶片,無 HBM 或特殊硬體
  • 60 天從軟體模型至客製 ASIC
  • 支援 Llama 3.1 8B 的 LoRA 微調
  • 本春推出更大推理模型,本冬推出前沿 LLM

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 7 個來源。

🔑 增強重點摘要

  • Taalas' HC1 chip hardwires Meta’s Llama 3.1 8B model directly into silicon using TSMC N6 (6nm) process, achieving over 16,000-17,000 tokens/second per user with under 1ms latency[1][2].
  • HC1 features a 815 mm² die size, ~250W power consumption, air-cooled compatibility, and uses on-chip SRAM for KV cache and fine-tuned weights, deployed as a PCIe card[1].
  • Turnaround time from a new model to working PCIe cards is approximately two months via a foundry-optimal workflow with TSMC[1].
  • Toronto-based Taalas raised $169M in recent funding, bringing total funding significantly beyond the initial $30M, with a team of 24 engineers targeting low-latency AI inference[6].
  • HC1 outperforms competitors like Nvidia, Cerebras, and Groq in tokens/second per user on Llama 3.1 8B, specialized for high-speed, low-latency inference without HBM[2].
📊 競品分析▸ Show
FeatureTaalas HC1Nvidia (implied)Cerebras (implied)Groq (implied)
Tokens/s per user (Llama 3.1 8B)>16,000 [2]Multiples slower [2]Multiples slower [2]Multiples slower [2]
Latency<1ms [1][2]Higher [2]Higher [2]Higher [2]
MemoryOn-chip SRAM, no HBM [1]HBM requiredSpecializedSpecialized
Power (single card)~250W [1]HigherHigherHigher

🛠️ 技術深入

  • Process/Fab: TSMC N6 (6nm)[1].
  • Die size: 815 mm²[1].
  • Form factor: PCIe card[1].
  • Power: ~250W per card, enabling 10-card server at ~2.5kW with standard air-cooling[1].
  • Memory: On-chip SRAM for KV cache and fine-tuned weights; no HBM or exotic hardware[1].
  • Model: Hardwired Llama 3.1 8B, supports LoRA fine-tuning[1][2].
  • Workflow: Foundry-optimal with TSMC for ~2-month model-to-PCIe turnaround[1].

🔮 前景展望AI analysis grounded in cited sources

Taalas' model-on-silicon approach could accelerate low-latency AI inference for edge and per-user applications, reducing reliance on general-purpose GPUs like Nvidia's and enabling cheaper, specialized hardware deployments, though limited to single-model runs per chip[1][2][5].

📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。