Zhuoyu Launches Physical AI Multimodal Model
Physical AI survival shift: Zhuoyu's multimodal model + new biz models for AV/robotics scale.
30-Second TL;DR
What Changed
Native multimodal model pre-trains vision, audio, actions together without language translation.
Why It Matters
This signals a paradigm shift in autonomous driving from expert models to scalable foundation models, potentially standardizing physical AI across mobility platforms. Zhuoyu's distribution strategies could accelerate adoption in L4 robotics, challenging incumbents.
What To Do Next
Integrate Zhuoyu's mobile AI SDK into your robotics prototype for quick physical AI testing.
Key Points
- •Native multimodal model pre-trains vision, audio, actions together without language translation.
- •Data mix: 30% vehicle, 30% robot, 40% internet first-person mobile videos.
- •Evolved from VLA 1.0 to VLA 2.0 paradigm, achieving ~70% zero-shot capability.
- •New business models: SDK distribution, open-source for partners, cloud-based action tokens.
- •Targets scale across vehicles, Robotaxi, RoboVan beyond traditional hardware sales.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.