💰TechCrunch AI•較早收集於 9m
Google 的 Genie 世界模型現在可模擬真實 Street View

💡利用 Google 龐大的 Street View 數據集,為機器人解鎖真實環境訓練能力。
⚡ 30-Second TL;DR
有什麼變化
將 Street View 整合至 Project Genie
為什麼重要
這項進展顯著降低了在真實且多樣化的環境中訓練機器人與自動化系統的門檻。
下一步行動
查閱 Project Genie 文件,評估您的機器人模擬流程是否能受益於真實的 Street View 數據。
誰應關注:Researchers & Academics
關鍵要點
- •將 Street View 整合至 Project Genie
- •實現互動式、可探索的世界模擬
- •支援動態天氣與罕見場景測試
🧠 深度解析
Web-grounded analysis with 21 cited sources.
🔑 增強重點摘要
- •Project Genie is powered by Genie 3, an 11-billion-parameter autoregressive transformer model, capable of generating real-time navigable 3D environments at 720p resolution and 24 frames per second.
- •The integration allows users to ground AI-generated worlds in real-world locations from Google Maps' dataset of 280 billion Street View images, applying various stylistic transformations like 'Ocean World' or 'Desert Sands'.
- •Genie 3 is designed to understand underlying physics, object permanence, and spatial consistency, which is crucial for creating realistic simulations and training embodied AI agents.
- •Project Genie is an experimental research prototype that was initially released to Google AI Ultra subscribers in the United States on January 29, 2026, and is now rolling out globally.
- •The model learns fine-grained controls and world dynamics from large datasets of unlabeled internet videos, including gameplay footage, without requiring explicit action labels.
📊 競品分析▸ Show
| Feature/Aspect | Google DeepMind Project Genie (Genie 3) | Odyssey AI Agora-1 | Odyssey AI Starchild-1 | OpenAI Sora / Google Veo |
|---|---|---|---|---|
| Core Function | Interactive, explorable 3D world simulation from text/images, real-world grounding | Multi-agent interactive 3D game simulation (e.g., GoldenEye environment) | Single-user interactive audio-video world model with text input | High-quality, passive video generation from text/images |
| Interactivity | Real-time, single-user navigation and dynamic environment changes | Real-time, up to four players interacting simultaneously in a shared world | Real-time, single-user interaction with synchronized visuals and sound | Pre-rendered video clips, no real-time interaction during playback |
| Resolution/FPS | 720p at 24 fps | Not explicitly stated for rendering, but simulates game state and renders individual perspectives in real-time | Up to 24 fps | High-quality video (specific resolution/FPS varies by model) |
| Real-world Data | Integrates Google Street View imagery for real-world location grounding | Focus on game environments, not explicitly real-world mapping | Not specified for real-world grounding | Not primarily focused on real-world grounding for interactive simulation |
| Primary Use Cases | AI agent training, robotics simulation, rapid game prototyping, creative content generation, education | Collaborative robotics, multi-agent AI training, emergent gameplay | AI agent training, multimodal interaction research | Content creation, visual storytelling |
| Model Type | 11-billion-parameter autoregressive transformer | Separates simulation (world state) and rendering (diffusion-based) | Interactive audio-video world model | Diffusion models (e.g., Sora) |
| Availability | Google AI Ultra subscribers (US, now global) | Early research preview on Odyssey website | Early research preview on Odyssey website | Varies (e.g., Sora in research preview, Veo more broadly available) |
🛠️ 技術深入
- Model Architecture: Genie 3 is an 11-billion-parameter autoregressive transformer, specifically adapted for visual sequence modeling.
- Operational Domain: It operates exclusively in the visual domain, generating pixel-based observations that users and AI agents can perceive and interact with.
- Frame Generation: The system generates each frame by considering the complete history of previously generated frames and the user's latest actions, ensuring consistency and coherence across extended sequences.
- Internal Mechanisms: It employs a visual tokenizer to compress frames into a latent space, a dynamics model to learn how these latent states evolve over time (predicting the next state given the current state and an action), and an action interface to map human inputs to the model's action tokens.
- Memory Architecture: Genie 3 incorporates a sophisticated memory system, including a short-term buffer (1-2 seconds for immediate consistency), a medium-term cache (10-30 seconds for recent interaction history), a long-term store (up to 1 minute for extended visual memory), and a semantic layer for high-level scene understanding and object relationships.
- Training Data & Methodology: The model is trained on large, diverse datasets of unlabeled internet videos, including footage of 2D platformer games and robotics. It learns to infer fine-grained controls and world dynamics without explicit action labels by predicting next frames in latent space and inferring actions that caused observed changes.
- Physics Simulation: While it does not implement explicit physics engines, Genie 3 learns physics patterns from its training data, allowing it to understand how objects should behave (e.g., water flow, object buoyancy, light behavior).
🔮 前景展望AI analysis grounded in cited sources
Project Genie will significantly accelerate AI agent training and robotics development.
It provides an unlimited curriculum of diverse, interactive, and physically consistent simulated environments, allowing AI agents and robots to learn through trial and error without real-world constraints or dangers.
The technology will democratize interactive content creation and virtual experience design.
By enabling the generation of complex 3D environments from simple text prompts or images, it lowers the barrier to entry for game development, creative exploration, and educational simulations.
Integration with Street View will lead to more personalized and realistic virtual tourism and urban planning tools.
Users can explore real-world locations with imaginative twists and dynamic changes, offering novel applications for virtual travel, architectural visualization, and scenario testing in urban environments.
⏳ 時間線
2001-XX
Stanford CityBlock Project, the inception of Google Street View technology.
2007-05
Google Street View officially launched in several U.S. cities.
2024-02
Genie 1 introduced, capable of generating 2D interactive environments from unlabeled internet videos.
2024-12
Genie 2 released, expanding capabilities to generate 3D environments with improved consistency.
2025-08
Genie 3 introduced, featuring higher-resolution world generations, increased memory, and real-time interaction.
2026-01-29
Project Genie, powered by Genie 3, released to Google AI Ultra subscribers in the United States.
2026-02
Waymo adopted Genie 3 to create a specialized 'Waymo World Model' for autonomous driving simulation.
2026-05-19
Project Genie integrates Google Street View data and rolls out globally to Google AI Ultra subscribers.
📎 來源 (21)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: TechCrunch AI ↗
