💰Stalecollected in 18m

Google Launches Gemini 3.5 Flash and Omni Models

Google Launches Gemini 3.5 Flash and Omni Models
PostLinkedIn
💰Read original on 钛媒体

💡Google's new agent-focused models are here, but the 5x price hike on the Flash tier is a critical cost factor.

⚡ 30-Second TL;DR

What Changed

Google I/O introduces Gemini Omni and 3.5 Flash models

Why It Matters

The price hike for the Flash model may force developers to re-evaluate their inference cost structures for high-volume agentic workflows.

What To Do Next

Benchmark your current agentic workflows against Gemini 3.5 Flash to determine if the performance gains justify the 5x cost increase.

Who should care:Developers & AI Engineers

Key Points

  • Google I/O introduces Gemini Omni and 3.5 Flash models
  • New models focus on enabling intelligent agent capabilities
  • Gemini 3.5 Flash pricing is 5x higher than the previous generation

🧠 Deep Insight

Web-grounded analysis with 27 cited sources.

🔑 Enhanced Key Takeaways

  • Gemini 3.5 Flash is positioned as Google's strongest agentic and coding model to date, demonstrating superior performance over Gemini 3.1 Pro on benchmarks such as Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%), while operating up to four times faster than comparable frontier models.
  • Gemini Omni represents a significant leap in multimodal AI, specifically for video generation and editing, allowing users to create and conversationally modify high-fidelity video content from diverse inputs including text, audio, images, and existing video, with an emphasis on consistent physics and character behavior.
  • The announced 5x cost increase for Gemini 3.5 Flash is relative to its predecessors, Gemini 3 Flash Preview and Gemini 3.1 Flash-Lite, with its per-token pricing set at $1.50 per million input tokens and $9.00 per million output tokens, making it approximately 40% cheaper than Gemini 3.1 Pro; however, its higher token consumption for agentic tasks can lead to overall higher operational costs.
  • Google is strategically advancing an 'Agentic Enterprise' vision, integrating Gemini 3.5 Flash and Omni as core components alongside new platforms like Antigravity 2.0 (an agent-first development platform) and Gemini Spark (a 24/7 personal AI agent designed for autonomous task execution).
  • Gemini 3.5 Flash is now the default model powering AI Mode in Google Search and the Gemini app, and it is broadly available across various Google products and developer platforms, including Google AI Studio and Antigravity.
📊 Competitor Analysis▸ Show
Feature/ModelGemini 3.5 FlashGemini 3.1 ProOpenAI GPT-5.5Anthropic Claude Opus 4.7xAI Grok 4.1
Input Pricing (per 1M tokens)$1.50$2.00$1.75 - $5.00 (varies by version/source)$3.00 - $5.00 (varies by version/source)$0.20
Output Pricing (per 1M tokens)$9.00$12.00$14.00 - $30.00 (varies by version/source)$15.00 - $25.00 (varies by version/source)$0.50
Key StrengthsAgentic tasks, coding, speed (4x faster than comparable frontier models), multi-step workflowsAcademic reasoning, long-document retrieval, deep knowledgeMathematical reasoning (e.g., AIME 2025), abstract reasoning (ARC-AGI-2)Software engineering (SWE-bench Verified), complex analysis, strong privacyCost-effectiveness, speed
Context Window1M tokens input, 64K tokens output1M tokens (up to 2M experimentally)Not explicitly stated for 5.5, but GPT-4o has 128K200K tokens (Opus 4.5)Not explicitly stated
MultimodalityText, image, audio, video input; text outputNatively multimodal (text, images, audio, video)Text, images, audioText, images, audio, videoNot explicitly stated

🛠️ Technical Deep Dive

  • Gemini 3.5 Flash is built upon the Gemini 3 Flash reasoning foundation, incorporating 'thinking levels' to manage the balance between quality, cost, and latency.
  • It supports a substantial 1 million token input context window and can generate up to 64,000 output tokens.
  • The model is natively multimodal, capable of processing text, images, audio, and video files as input, with its primary output being text.
  • Engineered by Google DeepMind, Gemini 3.5 Flash leverages Google's purpose-built AI infrastructure, specifically Tensor Processing Units (TPUs), for efficient training and deeper reasoning capabilities.
  • Key capabilities include function calling, structured output generation, search-as-a-tool, and code execution. It also features 'thinking' capabilities, which preserve encrypted reasoning context across API calls.
  • Gemini 3.5 Flash is optimized for agentic workflows, facilitating sub-agent deployment and multi-step tool use, and is designed to work with the Antigravity harness for orchestrating collaborative subagents.
  • Gemini Omni, the new video model, was trained on a diverse dataset comprising audio, video, image, and text data, with detailed text annotations for audio and video.
  • Omni is designed to simulate real-world physics and understand contextual information (e.g., historical facts) to generate more accurate and realistic video content.
  • All videos produced by Gemini Omni will be embedded with Google's SynthID digital watermark for content identification.

🔮 Future ImplicationsAI analysis grounded in cited sources

The emphasis on agentic capabilities will accelerate the development and adoption of autonomous AI systems across enterprises.
Gemini 3.5 Flash is specifically designed for long-horizon agentic tasks and multi-step workflows, supported by platforms like Antigravity 2.0 and Gemini Spark, enabling AI to take proactive actions and automate complex business processes.
Multimodal video generation and editing will become a new frontier for creative industries and content production.
Gemini Omni's ability to generate and conversationally edit high-fidelity video from diverse inputs, while maintaining physical consistency and offering avatar creation, opens up significant new possibilities for media creation and personalized content.
The rising cost of advanced AI models, despite performance gains, will drive a focus on efficiency and strategic model selection for specific use cases.
While Gemini 3.5 Flash offers strong performance, its higher per-token cost and increased token consumption for agentic tasks mean that total operational costs can exceed previous Flash models or even some Pro models, necessitating careful budgeting and workload optimization.

Timeline

2023-12
Google launches the Gemini model family (Nano, Pro), integrating it into the Bard chatbot and Pixel 8 Pro smartphone.
2024-02
The Bard chatbot is officially rebranded as Gemini.
2024-05
Gemini 1.5 Flash is announced at Google I/O 2024.
2025-01
Gemini 2.0 Flash is released as the new default model.
2026-02
Gemini 3.1 Pro is released in preview, characterized as a step forward in core reasoning.
2026-05
Google I/O 2026 introduces the Gemini 3.5 Flash and Gemini Omni models, emphasizing agentic capabilities and multimodal creation.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 钛媒体