SourceStalecollected in 14m

Zhuoyu Launches Physical AI Multimodal Model

Read original on 36氪
#embodied-ai#multimodal#autonomous-driving#foundation-model

Physical AI survival shift: Zhuoyu's multimodal model + new biz models for AV/robotics scale.

30-Second TL;DR

What Changed

Native multimodal model pre-trains vision, audio, actions together without language translation.

Why It Matters

This signals a paradigm shift in autonomous driving from expert models to scalable foundation models, potentially standardizing physical AI across mobility platforms. Zhuoyu's distribution strategies could accelerate adoption in L4 robotics, challenging incumbents.

What To Do Next

Integrate Zhuoyu's mobile AI SDK into your robotics prototype for quick physical AI testing.

Who should care:Developers & AI Engineers

Key Points

  • Native multimodal model pre-trains vision, audio, actions together without language translation.
  • Data mix: 30% vehicle, 30% robot, 40% internet first-person mobile videos.
  • Evolved from VLA 1.0 to VLA 2.0 paradigm, achieving ~70% zero-shot capability.
  • New business models: SDK distribution, open-source for partners, cloud-based action tokens.
  • Targets scale across vehicles, Robotaxi, RoboVan beyond traditional hardware sales.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.