⚛️較早收集於 74m

24人團隊硬剛英偉達!AMD前高管夢之隊新晶片每秒17000 token

24人團隊硬剛英偉達!AMD前高管夢之隊新晶片每秒17000 token
PostLinkedIn
⚛️閱讀原文: 量子位

💡17k tokens/sec at 1/10 Nvidia cost: potential inference revolution for AI devs.

⚡ 30-Second TL;DR

有什麼變化

24人團隊由前AMD高管組成

為什麼重要

此低成本高速度晶片可能顛覆英偉達AI硬體壟斷,讓新創與企業能更廉價部署大規模LLM。

下一步行動

Benchmark this chip against Nvidia H100 for your LLM inference workloads to assess cost savings.

誰應關注:Developers & AI Engineers

關鍵要點

  • 24人團隊由前AMD高管組成
  • 每秒17000 token推理速度
  • 成本僅英偉達的1/10
  • 直指英偉達市場霸主地位

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 6 個來源。

🔑 增強重點摘要

  • Taalas, led by a 24-person team of former AMD executives, unveiled the HC1 chip achieving 17,000 tokens per second per user on Llama 3.1 8B inference[1][2][3]
  • HC1 delivers ~73x higher throughput than Nvidia H200 and multiples above Cerebras (~2,000 tokens/sec) and Groq (~600 tokens/sec) on the same model[1][2][3]
  • Chip costs 1/10th the power of Nvidia equivalents and ~20x less to build, using air-cooled PCIe form factor[1][2][6]
  • Taalas raised $169 million in funding to develop model-specific AI chips challenging Nvidia dominance[1][5]
  • HC1 hardwires the entire model including weights onto the chip using mask ROM recall fabric, eliminating HBM and memory-compute bottlenecks[2][5]
📊 競品分析▸ Show
FeatureTaalas HC1Nvidia H200CerebrasGroq
Tokens/sec (Llama3.1-8B per user)17,000 [1][2][3]~230 (17k/73x) [1]~2,000 [2]~600 [2]
Power Consumption1/10th of Nvidia [1][5]Baseline [1]Not specified [2]Not specified [2]
Cost to Build20x less than SOTA [6]Baseline [6]Not specifiedNot specified
Form FactorPCIe card, ~250W air-cooled [2]GPU with HBM [5]Not specifiedNot specified

🛠️ 技術深入

  • Process/Fab: TSMC N6 (6nm)[2]
  • Die size: 815 mm²[2]
  • Power: ~250W per card; 10-card server ~2.5kW, air-cooled[2]
  • Architecture: Hardwires entire model (weights via mask ROM recall fabric), SRAM for KV cache and fine-tuned weights; single transistor per 4-bit module for matrix multiplications[2][5]
  • Memory: Eliminates HBM by merging storage and computation, no high-speed I/O or advanced packaging needed[2][5]
  • Form factor: PCIe card optimized for Llama 3.1 8B[2][5]

🔮 前景展望AI analysis grounded in cited sources

Taalas HC1 enables interactive frontier models with agentic behavior, reducing task times from hours to minutes at lower cost; unlocks new use cases like real-time reasoning with larger budgets for higher accuracy via multiple sampling or longer traces[2]. Model-specific chips challenge Nvidia by improving efficiency through specialization, potentially accelerating ubiquitous AI deployment[5][6].

時間線

2026-02
Taalas unveils HC1 chip with 17K tokens/sec on Llama 3.1 8B and raises $169M funding
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 量子位

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。