Google Introduces Gemini 3.8 Live Models

A new Gemini Live release may change real-time model choices for builders.
30-Second TL;DR
What Changed
Gemini 3.8 Live was introduced by Google DeepMind.
Why It Matters
If broadly available, the release could give developers another option for real-time multimodal or conversational applications. Practitioners should verify official documentation before committing to migration or production use.
What To Do Next
Check the official Gemini 3.8 Live documentation for API access, quotas, latency, and Extended Thinking controls before testing it.
Key Points
- •Gemini 3.8 Live was introduced by Google DeepMind.
- •A Gemini 3.8 Live Extended Thinking variant was also announced.
- •The excerpt does not provide benchmark results.
- •Pricing, access, and API availability are not specified in the supplied text.
Deep Insight
Background and context from public sources — not the original article. 9 sources cited.
Enhanced Key Takeaways
- •The models are built directly on top of the Gemini 3 Pro foundation architecture as native end-to-end speech-to-speech systems.
- •Gemini 3.8 Live models feature a 128K token (131,072) input context window and up to 64K tokens (65,536) of multimodal output generation.
- •The architecture supports asynchronous tool calling and parallel execution, allowing background API actions without interrupting active speech.
- •Gemini 3.8 Live Extended Thinking scored 82.6 on the Artificial Analysis Quality Index and 35.1 on the tau-banking agentic task completion benchmark.
- •Real-time audio usage is priced at $0.005 per minute for input audio and $0.018 per minute for output audio streaming.
Technical Deep Dive
- Foundational Architecture: Built on top of the Gemini 3 Pro base model using an end-to-end native audio approach rather than a cascaded ASR -> LLM -> TTS pipeline.
- Multimodal Context Window: Supports up to 131,072 tokens (128K) input across text, audio, images, and video, and generates up to 65,536 tokens (64K) in streaming text and audio output.
- Asynchronous Tool Calling: Features parallel reasoning capabilities enabling background API queries and function calls concurrently with continuous voice dialogue.
- Multilingual Support: Natively processes and dynamically auto-detects 97 languages with support for mid-sentence code-switching.
- Paralinguistic Preservation: Directly models and retains voice attributes including pitch, tone, hesitations, and inflection without text token intermediate bottlenecks.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-09Google DeepMind introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking models
Sources (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: DeepMind Blog ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.