⚛️量子位•Stalecollected in 2h
Nvidia's Jim Fan: VLA, Teleop Dead

💡Nvidia robotics boss kills VLA & teleop—new robot AI era incoming?
⚡ 30-Second TL;DR
What Changed
Jim Fan proclaims VLA models obsolete
Why It Matters
Signals potential paradigm shift in embodied AI, urging devs to explore post-VLA architectures.
What To Do Next
Review Jim Fan's full thesis on X and pivot your robot policy away from VLA.
Who should care:Developers & AI Engineers
Key Points
- •Jim Fan proclaims VLA models obsolete
- •Teleoperation declared dead in robotics
- •Statement from Nvidia's robotics head
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •Jim Fan advocates for a shift toward 'General Purpose World Models' that prioritize physical simulation and causal reasoning over the pattern-matching limitations inherent in current VLA architectures.
- •The critique of teleoperation centers on its inability to scale for long-horizon tasks, proposing instead that robots must learn from massive, diverse synthetic datasets generated within high-fidelity physics engines.
- •Nvidia's internal research direction is pivoting toward 'Embodied AI' that treats robotics as a control problem requiring world-model-based planning rather than just reactive vision-to-action mapping.
🛠️ Technical Deep Dive
- •VLA (Vision-Language-Action) models typically rely on autoregressive tokenization of visual inputs and motor commands, which Fan argues lacks true physical grounding.
- •Proposed alternative: World Models utilize latent dynamics models to predict future states (s_t+1) given an action (a_t), allowing for planning and counterfactual reasoning.
- •Shift toward 'Foundation Policies' trained on massive-scale simulation (e.g., Isaac Sim) to overcome the 'data bottleneck' of human-collected teleoperation data.
🔮 Future ImplicationsAI analysis grounded in cited sources
Teleoperation will be relegated to a niche data-collection method rather than a primary training paradigm.
The industry is hitting a scaling wall where human-in-the-loop data collection cannot keep pace with the requirements for general-purpose robot autonomy.
Robotics research will shift budget from hardware-heavy teleop setups to high-compute simulation clusters.
World model training requires massive synthetic data generation, necessitating a move toward compute-intensive simulation environments over physical data collection.
⏳ Timeline
2023-03
Jim Fan joins Nvidia as Lead of Embodied AI research.
2023-08
Nvidia introduces Voyager, an LLM-powered agent that plays Minecraft, signaling the shift toward autonomous agents.
2024-03
Nvidia announces Project GR00T, a foundation model for humanoid robots, emphasizing simulation-first training.
2025-01
Nvidia expands Isaac Lab, focusing on reinforcement learning and world model training for robotics.
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗
