Google Launches Gemini 3.5 Flash and Omni Models

💡Google's new agent-focused models are here, but the 5x price hike on the Flash tier is a critical cost factor.
⚡ 30-Second TL;DR
What Changed
Google I/O introduces Gemini Omni and 3.5 Flash models
Why It Matters
The price hike for the Flash model may force developers to re-evaluate their inference cost structures for high-volume agentic workflows.
What To Do Next
Benchmark your current agentic workflows against Gemini 3.5 Flash to determine if the performance gains justify the 5x cost increase.
Key Points
- •Google I/O introduces Gemini Omni and 3.5 Flash models
- •New models focus on enabling intelligent agent capabilities
- •Gemini 3.5 Flash pricing is 5x higher than the previous generation
🧠 Deep Insight
Web-grounded analysis with 27 cited sources.
🔑 Enhanced Key Takeaways
- •Gemini 3.5 Flash is positioned as Google's strongest agentic and coding model to date, demonstrating superior performance over Gemini 3.1 Pro on benchmarks such as Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%), while operating up to four times faster than comparable frontier models.
- •Gemini Omni represents a significant leap in multimodal AI, specifically for video generation and editing, allowing users to create and conversationally modify high-fidelity video content from diverse inputs including text, audio, images, and existing video, with an emphasis on consistent physics and character behavior.
- •The announced 5x cost increase for Gemini 3.5 Flash is relative to its predecessors, Gemini 3 Flash Preview and Gemini 3.1 Flash-Lite, with its per-token pricing set at $1.50 per million input tokens and $9.00 per million output tokens, making it approximately 40% cheaper than Gemini 3.1 Pro; however, its higher token consumption for agentic tasks can lead to overall higher operational costs.
- •Google is strategically advancing an 'Agentic Enterprise' vision, integrating Gemini 3.5 Flash and Omni as core components alongside new platforms like Antigravity 2.0 (an agent-first development platform) and Gemini Spark (a 24/7 personal AI agent designed for autonomous task execution).
- •Gemini 3.5 Flash is now the default model powering AI Mode in Google Search and the Gemini app, and it is broadly available across various Google products and developer platforms, including Google AI Studio and Antigravity.
📊 Competitor Analysis▸ Show
| Feature/Model | Gemini 3.5 Flash | Gemini 3.1 Pro | OpenAI GPT-5.5 | Anthropic Claude Opus 4.7 | xAI Grok 4.1 |
|---|---|---|---|---|---|
| Input Pricing (per 1M tokens) | $1.50 | $2.00 | $1.75 - $5.00 (varies by version/source) | $3.00 - $5.00 (varies by version/source) | $0.20 |
| Output Pricing (per 1M tokens) | $9.00 | $12.00 | $14.00 - $30.00 (varies by version/source) | $15.00 - $25.00 (varies by version/source) | $0.50 |
| Key Strengths | Agentic tasks, coding, speed (4x faster than comparable frontier models), multi-step workflows | Academic reasoning, long-document retrieval, deep knowledge | Mathematical reasoning (e.g., AIME 2025), abstract reasoning (ARC-AGI-2) | Software engineering (SWE-bench Verified), complex analysis, strong privacy | Cost-effectiveness, speed |
| Context Window | 1M tokens input, 64K tokens output | 1M tokens (up to 2M experimentally) | Not explicitly stated for 5.5, but GPT-4o has 128K | 200K tokens (Opus 4.5) | Not explicitly stated |
| Multimodality | Text, image, audio, video input; text output | Natively multimodal (text, images, audio, video) | Text, images, audio | Text, images, audio, video | Not explicitly stated |
🛠️ Technical Deep Dive
- Gemini 3.5 Flash is built upon the Gemini 3 Flash reasoning foundation, incorporating 'thinking levels' to manage the balance between quality, cost, and latency.
- It supports a substantial 1 million token input context window and can generate up to 64,000 output tokens.
- The model is natively multimodal, capable of processing text, images, audio, and video files as input, with its primary output being text.
- Engineered by Google DeepMind, Gemini 3.5 Flash leverages Google's purpose-built AI infrastructure, specifically Tensor Processing Units (TPUs), for efficient training and deeper reasoning capabilities.
- Key capabilities include function calling, structured output generation, search-as-a-tool, and code execution. It also features 'thinking' capabilities, which preserve encrypted reasoning context across API calls.
- Gemini 3.5 Flash is optimized for agentic workflows, facilitating sub-agent deployment and multi-step tool use, and is designed to work with the Antigravity harness for orchestrating collaborative subagents.
- Gemini Omni, the new video model, was trained on a diverse dataset comprising audio, video, image, and text data, with detailed text annotations for audio and video.
- Omni is designed to simulate real-world physics and understand contextual information (e.g., historical facts) to generate more accurate and realistic video content.
- All videos produced by Gemini Omni will be embedded with Google's SynthID digital watermark for content identification.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (27)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- google.com
- llm-stats.com
- datacamp.com
- indiatimes.com
- siliconangle.com
- macrumors.com
- mashable.com
- knightli.com
- trendingtopics.eu
- buildfastwithai.com
- simonwillison.net
- the-decoder.com
- cnet.com
- youtube.com
- engadget.com
- latent.space
- seroundtable.com
- intuitionlabs.ai
- truinc.com
- medium.com
- deepmind.google
- gradually.ai
- fivetran.com
- llm-stats.com
- deepmind.google
- google.dev
- google.dev
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 钛媒体 ↗


