🐼Stalecollected in 29m

XPeng 2nd-Gen VLA Rolls Out This Month

XPeng 2nd-Gen VLA Rolls Out This Month
PostLinkedIn
🐼Read original on Pandaily

💡XPeng's 2nd-gen VLA launches—embodied AI boost for AV builders

⚡ 30-Second TL;DR

What Changed

Second-gen VLA model launches this month

Why It Matters

Accelerates XPeng's embodied AI for autonomous driving. Competitive X9 pricing strengthens EV market position against rivals like Tesla.

What To Do Next

Benchmark XPeng's new VLA against open-source VLAs for AV simulation tasks.

Who should care:Developers & AI Engineers

Key Points

  • Second-gen VLA model launches this month
  • 2026 X9 BEV priced from $42,800 to $51,000
  • VLA enables vision-language-action integration for vehicles
  • Part of XPeng's AI-driven EV advancements

🧠 Deep Insight

Web-grounded analysis with 8 cited sources.

🔑 Enhanced Key Takeaways

  • XPeng VLA 2.0 adopts a 'Vision-Implicit Token-Action' architecture, bypassing traditional language translation for direct visual-to-action generation.[1][2][3]
  • The model was trained on nearly 100 million unannotated driving video clips, equivalent to 65,000 years of human driving experience.[3][4][6]
  • Powered by Turing AI chips delivering 2,250 TOPS per chip, enabling 12x inference efficiency gains and deployment across cars, robots, and flying cars.[1][2][6]
  • VLA 2.0 serves as XPeng's first mass-produced physical world model, capable of self-evolving learning and cross-domain applications.[3][5]

🛠️ Technical Deep Dive

  • Architecture shift to 'Vision-Implicit Token-Action' (V-Implicit Token-A) eliminates V-L-A language bottleneck, enabling end-to-end V-A output for human-like reflexes.[1][2][3]
  • Trained without data annotation using ~100 million real-world driving clips covering long-tail scenarios; generates adversarial scenarios for self-improvement.[3][4][6]
  • Deployed via custom compiler on Turing AI chip (2,250 TOPS/chip); model parameters 10x larger than mainstream, with 12x inference efficiency.[1][6]
  • Combines with VLT (Vision-Language-Thinking) and VLM for high-order capabilities like conversation, walking, and interaction in robots like IRON.[2][3][6]
  • XPeng World Base Models trained at 1B, 3B, 7B, 72B parameters; four-stage factory process: pre-training, distillation, continued training, vehicle deployment.[7]

🔮 Future ImplicationsAI analysis grounded in cited sources

XPeng VLA 2.0 will enable Level 4 autonomy in consumer EVs by mid-2026
FastDriveVLA optimization on VLA 2.0 architecture, trained on massive data, boosts narrow-road performance 13x and fits within vehicle power constraints.[4][6]
Cross-domain VLA deployment will accelerate XPeng humanoid robot mass production by end-2026
VLA 2.0 powers IRON robot with VLT+VLA+VLM stack on three Turing chips, supporting real-time physical tasks amid Guangzhou data factory plans.[2][3][6]
VLA revenue generation via SDK licensing will exceed vehicle sales contributions within 10 years
He Xiaopeng states second-gen VLA enables third-party apps like AR modules, with open-sourcing for commercial partners.[5][8]

Timeline

2024-Q2
Turing AI chip enters mass production for VLA deployment.
2025-11
XPeng AI Day unveils VLA 2.0, Robotaxi, next-gen IRON robot, and flying car applications.
2025-11
FastDriveVLA optimization introduced for L4 autonomy on VLA 2.0 architecture.
2025-12
Detailed analysis of FastDriveVLA confirms 13x narrow-road performance gains.
2026-03
Second-gen VLA rolls out for vehicles including 2026 X9 BEV.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily