來源DeepMind Blog•近期收集於 5h
Google 推出 Gemini 3.8 Live 模型

#real-time-ai#extended-thinking#model-releasegemini-3.8-livegoogle-deepmindgemini-3.8-live
新的 Gemini Live 發布可能改變建構者對即時模型的選擇。
30 秒速覽
有什麼變化
Google DeepMind 推出了 Gemini 3.8 Live。
為什麼重要
若廣泛提供,這項發布可能為即時多模態或對話式應用提供另一個選擇。實務人員在決定遷移或投入正式環境前,應先確認官方文件。
下一步行動
在測試前,查閱官方 Gemini 3.8 Live 文件,確認 API 存取、配額、延遲與 Extended Thinking 控制項。
誰應關注:Developers & AI Engineers
關鍵要點
- •Google DeepMind 推出了 Gemini 3.8 Live。
- •公司也宣布 Gemini 3.8 Live Extended Thinking 版本。
- •摘錄未提供基準測試結果。
- •提供的文字未說明定價、存取方式與 API 可用性。
深度解析
背景與延伸:來自公開資料,非原文內容。引用 9 個來源。
增強重點摘要
- •Gemini 3.8 Live 系列建立在 Gemini 3 Pro 基礎模型之上,採用端到端原生語音對語音架構,全面取代由 ASR、LLM 到 TTS 的傳統串聯管線,能保留音調、語氣和停頓等副語言線索。
- •模型具備 128K(131,072 tokens)多模態輸入上下文窗口,以及高達 64K(65,536 tokens)的多模態串流輸出能力,支援文字、音訊、圖片與視訊輸入。
- •支援原生辨識 97 種語言,並具備句中動態切換語言(mid-sentence switching)能力,同時支援非同步函式呼叫,能在對話期間於背景執行 API 調用而無需中斷發言。
- •Gemini 3.8 Live Extended Thinking 在 Artificial Analysis Quality Index 取得 82.6 分並登上語音對語音榜首,並在 tau-banking 代理任務基準獲得 35.1 分。
- •即時語音 API 定價為音訊輸入每分鐘 0.005 美元、音訊輸出每分鐘 0.018 美元,並全面整合至 Google AI Studio、Gemini Live API、Google Workspace 及 Gemini 消費端應用。
技術深入
- 基礎模型架構:以 Gemini 3 Pro 為基底構建,具備端到端原生語音對語音(Speech-to-Speech)推論能力,無需分段經過語音辨識(ASR)與文字轉語音(TTS)管道。
- 上下文長度規格:支援 131,072 tokens(128K)的多模態輸入(包含文字、音訊、圖像與視訊)以及 65,536 tokens(64K)的多模態(文字與即時串流語音)輸出。
- 非同步工具呼叫:支援平行推論與非同步函式調用(Asynchronous Function Calling),模型可在背景觸發外部 API 或檢索資料,避免語音對話產生不自然的中斷或延遲。
- 多語言辨識機制:原生內建 97 種語言自動偵測,並支援對話進行中流暢地在語句中切換語言(Mid-Sentence Switching)。
- 副語言特徵捕捉:直接由原生音訊端點學習並保留說話者的音高(pitch)、語調變化(inflection)、語氣(tone)及猶豫停頓(hesitations)。
前景展望基於引用來源的 AI 分析
串聯式語音代理架構將加速被原生端到端語音模型淘汰
Gemini 3.8 Live 證明原生語音模型可同時達成保留語氣細節與非同步背景任務處理,消除了傳統 ASR-LLM-TTS 串聯架構的高延遲與資訊丟失缺點。
即時語音 AI 代理的跨語言客服應用成本將大幅降低
以每分鐘 0.005 美元輸入與 0.018 美元輸出的低定價,加上 97 種語言動態切換能力,將推動跨國企業大規模部署語音代理。
時間線
2026-09
Google DeepMind 推出 Gemini 3.8 Live 與 Gemini 3.8 Live Extended Thinking 模型
- 2026-09Google DeepMind 推出 Gemini 3.8 Live 與 Gemini 3.8 Live Extended Thinking 模型
來源 (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
1Google Releases Gemini 3 8 Live and 3 8 Live Extended Thinking for Production Grade Voice Agentsmarktechpost.com2Build Real Time Voice Applications Gemini Audioblog.google3Gemini 3 8 Livejetstream.blog4Gemini 3 8 Live Gemini 3 8 Live Extended Thinkingblog.google5Gemini 3 8 Audiodeepmind.google6google.devai.google.dev7Google Launches Gemini 3 8 Live Audio Modelsandroidheadlines.com8Gemini 3 8 Live Thinks in the Background While You Talk and the Benchmarks Are Hard to Ignore Vqety7y4wdaily.dev93 8 Flash and 3 8 Flash Cyberblog.google
AI 週報
閱讀本週精選 AI 大事摘要 →
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: DeepMind Blog ↗
每週電子報
每週一封,可隨時退訂。