較早收集於 2h

Google I/O 2026:16 項 AI 產品更新與 Agent 戰略

Google I/O 2026:16 項 AI 產品更新與 Agent 戰略
PostLinkedIn
閱讀原文: 雷峰网

💡Google 的大規模 Agent 攻勢:16 項更新、每月 3 兆 token 處理量,以及全新的搜尋典範。

⚡ 30-Second TL;DR

有什麼變化

Gemini 3.5 Flash 提供 4 倍輸出速度,成本僅為前代模型的一半。

為什麼重要

Google 的戰略從「AI 作為聊天機器人」轉向「AI 作為持久的 Agent 層」,迫使競爭對手加速其 Agent 生態系統的整合。

下一步行動

評估使用 Gemini 3.5 Flash 處理高流量推理任務,以優化成本與延遲。

誰應關注:Enterprise & Security Teams

關鍵要點

  • Gemini 3.5 Flash 提供 4 倍輸出速度,成本僅為前代模型的一半。
  • 搜尋 AI 模式可針對複雜查詢動態生成互動式 UI 與儀表板。
  • Google 將 AI Agent 整合至 Workspace、Android 與硬體中,實現 24/7 主動式協助。

🧠 深度解析

Web-grounded analysis with 30 cited sources.

🔑 增強重點摘要

  • Gemini 3.5 Flash is specifically optimized for agentic workflows and coding tasks, demonstrating competitive performance against previous Pro models in these areas while offering increased speed and reduced cost.
  • The revamped Search AI Mode now enables users to create and manage persistent 'information agents' directly from the search interface, which can autonomously track ongoing tasks like stock monitoring or apartment listings and deliver proactive alerts.
  • Google introduced Gemini Spark, a cloud-based, 24/7 personal AI agent designed to proactively perform complex tasks such as drafting emails from recent threads, monitoring inboxes for specific content, or scanning financial statements for new subscriptions.
  • Google unveiled Gemini Omni, a new multimodal AI model initially focused on high-quality video generation and editing, with future plans to expand its capabilities to include image and text generation.
  • The company significantly expanded its 'Antigravity' platform for agent development, introducing Antigravity 2.0 as a standalone desktop application, a command-line interface (CLI), and an SDK, alongside 'Managed Agents' in the Gemini API for hosted agent runtimes.
📊 競品分析▸ Show
Feature/ModelGoogle Gemini 3.5 Flash (May '26)OpenAI GPT-4o (Nov '24)OpenAI GPT-4o Mini (Jul '24)Anthropic Claude Opus 4.7 (max)Anthropic Claude Haiku (Aug '24)
Input Cost (per 1M tokens)$1.50$2.50$0.15$5.00$1.00
Output Cost (per 1M tokens)$9.00$10.00$6.00$5.00$5.00
Throughput (tokens/second)~289~149.1N/A~50~93
Context Window1M tokens128K tokensN/A (large context)N/A (large context)N/A (large context)
Output Token Limit65,53616,000N/AN/AN/A
GPQA Benchmark92.2%54.3%N/AN/AN/A
Coding Index45.016.7GoodN/AN/A
Intelligence Index55.317.3N/A57.3 (top 3)N/A
Key StrengthsHigh speed, cost-efficient, strong for coding & agentic tasks, large output capacity.Maintains GPT-4 Turbo intelligence, multimodal input.Cost-efficient, good for chaining/parallel calls, real-time support.High quality, strong reasoning, long-form complex tasks.Incredible speed, cost, text processing, instruction following.

🛠️ 技術深入

  • Gemini 3.5 Flash Model Specifications:
    • Inputs: Multimodal, accepting text, images, audio, video, and PDFs.
    • Output: Primarily text-only.
    • Context Window: 1 million tokens, suitable for extensive documents and conversation history.
    • Output Token Limit: Up to 65,536 tokens, surpassing Gemini 3.1 Pro's 32,768 tokens, making it suitable for long generations.
    • Latency: Significantly lower than Pro models, optimized for rapid response in production environments.
    • Pricing: Substantially cheaper per million input/output tokens compared to Pro models, with cached input pricing at $0.15/1M tokens (90% cheaper).
    • Tooling: Supports function calling, structured output, code execution, and search-as-a-tool.
    • Architecture: Built on the Gemini 3 Flash reasoning foundation, incorporating explicit thinking levels to balance quality, cost, and latency.
    • Speed: Achieves approximately 289 output tokens per second.
  • Search AI Mode Implementation:
    • Utilizes a custom version of Gemini and a 'query fan-out' technique, which breaks down complex queries into subtopics and searches them simultaneously across multiple data sources.
    • Integrates 'Generative UI' capabilities, allowing the AI to interpret user intent and dynamically generate interactive user interfaces, tools, and simulations on the fly for complex queries.
    • Supports multimodal input, including text, voice, images, and even real-time screen content via Google Lens integration.
  • Antigravity Agent Development Platform:
    • Antigravity 2.0: A standalone desktop application designed for orchestrating multiple AI agents in parallel, including scheduled tasks for background automation.
    • Antigravity CLI: A terminal-first interface for developers to spin up agents without a graphical user interface.
    • Antigravity SDK: Provides programmatic access to the underlying agent harness, allowing developers to host agents on their chosen infrastructure.
    • Managed Agents in Gemini API: Enables developers to build, define, and run hosted agents on Google's infrastructure with a single Gemini API key, providing an isolated, ephemeral Linux environment for reasoning, code execution, and web browsing.
    • Unified Agent Harness: A single runtime layer that handles reasoning, tool calls, and code execution, deployed across various Google surfaces, connecting individual developer tools with enterprise platforms.

🔮 前景展望AI analysis grounded in cited sources

Google's AI agent strategy will fundamentally transform user interaction from explicit commands to proactive, autonomous assistance.
The introduction of Gemini Spark and information agents in Search indicates a shift towards AI systems that continuously monitor, manage, and act on behalf of users without constant prompting, moving beyond simple question-answering.
The deep integration of Gemini across Google's ecosystem, including Android and new smart glasses, will establish AI as a core operating system layer rather than merely an application.
Google's vision positions Gemini as infrastructure within Android 17 and upcoming smart glasses, enabling AI to observe, infer, warn, and act directly within the device's operating environment, making the phone less a set of apps and more a context-rich endpoint.
Google's dual-lane agent development strategy, offering both consumer-developer APIs and an enterprise platform, will attract a broad range of developers and accelerate the creation of diverse AI agent solutions.
By providing low-friction entry points for individual developers via Antigravity and a governed tier for enterprises through the Gemini Enterprise Agent Platform, Google aims to foster a wide and integrated ecosystem of agent builders.

時間線

2017
Google researchers present the transformer architecture, foundational to many LLMs.
2023-05-10
Google announces Gemini at I/O, positioned as a multimodal successor to PaLM 2.
2023-12-06
Gemini (Pro, Deep Think, Flash, Flash Lite) officially announced and released in beta.
2024-02
Bard and Duet AI unified under the Gemini brand; Gemini Advanced with Ultra 1.0 released.
2024-02
Gemini 1.5 released with a new architecture and 1-million-token context window.
2025-04-09
Google introduces the Agent Development Kit (ADK) and Agent2Agent Protocol (A2A) at Google Cloud NEXT 2025.
2025-11-18
Generative UI capabilities integrated into Google Search AI Mode.
2026-03-16
Google Search AI Mode Canvas, allowing creation of interactive tools and dashboards, becomes available in the U.S.
2026-05-19
Google I/O 2026 announcements, including Gemini 3.5 Flash, Gemini Omni, Search AI Mode overhaul, Gemini Spark, and Antigravity 2.0.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 雷峰网