SourceRecentcollected in 5h

Google Introduces Gemini 3.8 Live Models

Read original on DeepMind Blog
#real-time-ai#extended-thinking#model-release

A new Gemini Live release may change real-time model choices for builders.

30-Second TL;DR

What Changed

Gemini 3.8 Live was introduced by Google DeepMind.

Why It Matters

If broadly available, the release could give developers another option for real-time multimodal or conversational applications. Practitioners should verify official documentation before committing to migration or production use.

What To Do Next

Check the official Gemini 3.8 Live documentation for API access, quotas, latency, and Extended Thinking controls before testing it.

Who should care:Developers & AI Engineers

Key Points

  • Gemini 3.8 Live was introduced by Google DeepMind.
  • A Gemini 3.8 Live Extended Thinking variant was also announced.
  • The excerpt does not provide benchmark results.
  • Pricing, access, and API availability are not specified in the supplied text.
Key numbers$0.005$0.018

Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

Enhanced Key Takeaways

  • The models are built directly on top of the Gemini 3 Pro foundation architecture as native end-to-end speech-to-speech systems.
  • Gemini 3.8 Live models feature a 128K token (131,072) input context window and up to 64K tokens (65,536) of multimodal output generation.
  • The architecture supports asynchronous tool calling and parallel execution, allowing background API actions without interrupting active speech.
  • Gemini 3.8 Live Extended Thinking scored 82.6 on the Artificial Analysis Quality Index and 35.1 on the tau-banking agentic task completion benchmark.
  • Real-time audio usage is priced at $0.005 per minute for input audio and $0.018 per minute for output audio streaming.

Technical Deep Dive

  • Foundational Architecture: Built on top of the Gemini 3 Pro base model using an end-to-end native audio approach rather than a cascaded ASR -> LLM -> TTS pipeline.
  • Multimodal Context Window: Supports up to 131,072 tokens (128K) input across text, audio, images, and video, and generates up to 65,536 tokens (64K) in streaming text and audio output.
  • Asynchronous Tool Calling: Features parallel reasoning capabilities enabling background API queries and function calls concurrently with continuous voice dialogue.
  • Multilingual Support: Natively processes and dynamically auto-detects 97 languages with support for mid-sentence code-switching.
  • Paralinguistic Preservation: Directly models and retains voice attributes including pitch, tone, hesitations, and inflection without text token intermediate bottlenecks.

Future ImplicationsAI analysis grounded in cited sources

Voice agent architectures will transition entirely away from cascaded ASR-LLM-TTS pipelines
Native speech-to-speech models preserve paralinguistic nuances and drastically reduce dialogue latency compared to sequential multi-model stacks.
Real-time voice agents will increasingly manage complex business workflows via asynchronous execution
Allowing background function calling without interrupting spoken conversational flow enables agentic tasks like live account management and database lookups in uninterrupted dialogue.

Timeline

2026-09
Google DeepMind introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking models

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: DeepMind Blog

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.