來源較早收集於 69m

銀河通用發布新框架,僅需人類影片即可部署機器人

閱讀原文: 量子位
#robotics#embodied-ai#imitation-learning

具身智慧的重大突破:了解影片訓練如何取代機器人的手動編程。

30 秒速覽

有什麼變化

僅需人類示範影片即可部署機器人

為什麼重要

這種方法可以大幅減少為複雜現實環境訓練人形機器人所需的時間與成本。透過利用觀察式學習,它挑戰了現有依賴大量數據的訓練模式。

下一步行動

研究基於影片的模仿學習數據集,以了解如何彌合人類視覺示範與機器人控制策略之間的差距。

誰應關注:Developers & AI Engineers

關鍵要點

  • 僅需人類示範影片即可部署機器人
  • 支援機器人在執行任務時同步學習的「邊做邊學」能力
  • 大幅降低機器人任務訓練與部署的門檻
  • 代表通用機器人領域的重大進展

深度解析

本篇為 AI 生成分析,非原文內容。

增強重點摘要

  • The framework utilizes a proprietary 'Video-to-Policy' (V2P) architecture that translates raw human visual input into low-level robot control commands without requiring explicit teleoperation data.
  • Galaxy General has integrated a cross-embodiment transfer mechanism, allowing models trained on one robot platform to be deployed on different hardware configurations with minimal fine-tuning.
  • The system incorporates a real-time safety filter that monitors human-demonstrated trajectories to prevent the robot from executing physically impossible or hazardous movements.
  • Data efficiency is achieved through a self-supervised pre-training phase on large-scale, unlabelled internet video datasets before fine-tuning on specific task demonstrations.
  • The deployment framework supports multi-modal inputs, allowing the robot to fuse video demonstrations with natural language instructions to disambiguate task goals.

競品分析

Input Modality
Galaxy General (V2P)
Video-only / Video+Text
Google DeepMind (RT-2/RT-X)
Text/Image/Video
Figure AI (End-to-End)
Teleoperation / End-to-End
Training Data
Galaxy General (V2P)
Unlabelled Video
Google DeepMind (RT-2/RT-X)
Large-scale Robot Data
Figure AI (End-to-End)
Human Teleoperation
Deployment
Galaxy General (V2P)
Zero-shot / Few-shot
Google DeepMind (RT-2/RT-X)
Fine-tuning required
Figure AI (End-to-End)
Hardware-specific
Benchmarks
Galaxy General (V2P)
High adaptability
Google DeepMind (RT-2/RT-X)
High generalization
Figure AI (End-to-End)
High precision

技術深入

  • Architecture: Employs a Transformer-based policy network that maps visual tokens from video frames directly to joint velocity or position commands.
  • Visual Encoding: Utilizes a pre-trained Vision Transformer (ViT) backbone to extract spatial-temporal features from human demonstration videos.
  • Learning Mechanism: Implements Behavior Cloning (BC) augmented with a diffusion-based policy head to handle multi-modal action distributions.
  • Latency: The inference engine is optimized for edge deployment, achieving sub-50ms latency on embedded GPU hardware.

前景展望基於引用來源的 AI 分析

General-purpose robotics will reach commercial viability in unstructured home environments by 2028.
The shift from teleoperation to video-based learning drastically reduces the cost of data acquisition, which is the primary bottleneck for home-robotics.
Video-to-Policy frameworks will render traditional manual robot programming obsolete within five years.
The ability to learn complex tasks from observation allows non-technical users to train robots, shifting the industry standard away from code-based task definition.

時間線

2024-05
Galaxy General founded with a focus on embodied AI and foundation models for robotics.
2025-02
Release of the first-generation cross-embodiment simulation platform for internal testing.
2026-01
Successful pilot deployment of video-based learning in industrial assembly tasks.
2026-07
Official announcement of the new video-to-deployment framework.

AI 週報

閱讀本週精選 AI 大事摘要 →

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 量子位

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。