💰較早收集於 18m

Google 發布 Gemini 3.5 Flash 與 Omni 模型

Google 發布 Gemini 3.5 Flash 與 Omni 模型
PostLinkedIn
💰閱讀原文: 钛媒体

💡Google 推出專注於代理的新模型,但 Flash 版本 5 倍的價格漲幅是開發者必須關注的成本關鍵。

⚡ 30-Second TL;DR

有什麼變化

Google I/O 發布 Gemini Omni 與 3.5 Flash 模型

為什麼重要

Flash 模型價格的上漲可能會迫使開發者重新評估高流量代理工作流的推理成本結構。

下一步行動

對比您當前的代理工作流與 Gemini 3.5 Flash 的性能,以評估性能提升是否足以抵銷 5 倍的成本漲幅。

誰應關注:Developers & AI Engineers

關鍵要點

  • Google I/O 發布 Gemini Omni 與 3.5 Flash 模型
  • 新模型專注於賦能智能代理功能
  • Gemini 3.5 Flash 的定價較前代高出 5 倍

🧠 深度解析

Web-grounded analysis with 27 cited sources.

🔑 增強重點摘要

  • Gemini 3.5 Flash is positioned as Google's strongest agentic and coding model to date, demonstrating superior performance over Gemini 3.1 Pro on benchmarks such as Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%), while operating up to four times faster than comparable frontier models.
  • Gemini Omni represents a significant leap in multimodal AI, specifically for video generation and editing, allowing users to create and conversationally modify high-fidelity video content from diverse inputs including text, audio, images, and existing video, with an emphasis on consistent physics and character behavior.
  • The announced 5x cost increase for Gemini 3.5 Flash is relative to its predecessors, Gemini 3 Flash Preview and Gemini 3.1 Flash-Lite, with its per-token pricing set at $1.50 per million input tokens and $9.00 per million output tokens, making it approximately 40% cheaper than Gemini 3.1 Pro; however, its higher token consumption for agentic tasks can lead to overall higher operational costs.
  • Google is strategically advancing an 'Agentic Enterprise' vision, integrating Gemini 3.5 Flash and Omni as core components alongside new platforms like Antigravity 2.0 (an agent-first development platform) and Gemini Spark (a 24/7 personal AI agent designed for autonomous task execution).
  • Gemini 3.5 Flash is now the default model powering AI Mode in Google Search and the Gemini app, and it is broadly available across various Google products and developer platforms, including Google AI Studio and Antigravity.
📊 競品分析▸ Show
Feature/ModelGemini 3.5 FlashGemini 3.1 ProOpenAI GPT-5.5Anthropic Claude Opus 4.7xAI Grok 4.1
Input Pricing (per 1M tokens)$1.50$2.00$1.75 - $5.00 (varies by version/source)$3.00 - $5.00 (varies by version/source)$0.20
Output Pricing (per 1M tokens)$9.00$12.00$14.00 - $30.00 (varies by version/source)$15.00 - $25.00 (varies by version/source)$0.50
Key StrengthsAgentic tasks, coding, speed (4x faster than comparable frontier models), multi-step workflowsAcademic reasoning, long-document retrieval, deep knowledgeMathematical reasoning (e.g., AIME 2025), abstract reasoning (ARC-AGI-2)Software engineering (SWE-bench Verified), complex analysis, strong privacyCost-effectiveness, speed
Context Window1M tokens input, 64K tokens output1M tokens (up to 2M experimentally)Not explicitly stated for 5.5, but GPT-4o has 128K200K tokens (Opus 4.5)Not explicitly stated
MultimodalityText, image, audio, video input; text outputNatively multimodal (text, images, audio, video)Text, images, audioText, images, audio, videoNot explicitly stated

🛠️ 技術深入

  • Gemini 3.5 Flash is built upon the Gemini 3 Flash reasoning foundation, incorporating 'thinking levels' to manage the balance between quality, cost, and latency.
  • It supports a substantial 1 million token input context window and can generate up to 64,000 output tokens.
  • The model is natively multimodal, capable of processing text, images, audio, and video files as input, with its primary output being text.
  • Engineered by Google DeepMind, Gemini 3.5 Flash leverages Google's purpose-built AI infrastructure, specifically Tensor Processing Units (TPUs), for efficient training and deeper reasoning capabilities.
  • Key capabilities include function calling, structured output generation, search-as-a-tool, and code execution. It also features 'thinking' capabilities, which preserve encrypted reasoning context across API calls.
  • Gemini 3.5 Flash is optimized for agentic workflows, facilitating sub-agent deployment and multi-step tool use, and is designed to work with the Antigravity harness for orchestrating collaborative subagents.
  • Gemini Omni, the new video model, was trained on a diverse dataset comprising audio, video, image, and text data, with detailed text annotations for audio and video.
  • Omni is designed to simulate real-world physics and understand contextual information (e.g., historical facts) to generate more accurate and realistic video content.
  • All videos produced by Gemini Omni will be embedded with Google's SynthID digital watermark for content identification.

🔮 前景展望AI analysis grounded in cited sources

The emphasis on agentic capabilities will accelerate the development and adoption of autonomous AI systems across enterprises.
Gemini 3.5 Flash is specifically designed for long-horizon agentic tasks and multi-step workflows, supported by platforms like Antigravity 2.0 and Gemini Spark, enabling AI to take proactive actions and automate complex business processes.
Multimodal video generation and editing will become a new frontier for creative industries and content production.
Gemini Omni's ability to generate and conversationally edit high-fidelity video from diverse inputs, while maintaining physical consistency and offering avatar creation, opens up significant new possibilities for media creation and personalized content.
The rising cost of advanced AI models, despite performance gains, will drive a focus on efficiency and strategic model selection for specific use cases.
While Gemini 3.5 Flash offers strong performance, its higher per-token cost and increased token consumption for agentic tasks mean that total operational costs can exceed previous Flash models or even some Pro models, necessitating careful budgeting and workload optimization.

時間線

2023-12
Google launches the Gemini model family (Nano, Pro), integrating it into the Bard chatbot and Pixel 8 Pro smartphone.
2024-02
The Bard chatbot is officially rebranded as Gemini.
2024-05
Gemini 1.5 Flash is announced at Google I/O 2024.
2025-01
Gemini 2.0 Flash is released as the new default model.
2026-02
Gemini 3.1 Pro is released in preview, characterized as a step forward in core reasoning.
2026-05
Google I/O 2026 introduces the Gemini 3.5 Flash and Gemini Omni models, emphasizing agentic capabilities and multimodal creation.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 钛媒体