Fudan-linked team unveils robot-native world action model

💡New embodied AI architecture from a high-growth startup team; potential breakthrough in robotic world modeling.
⚡ 30-Second TL;DR
What Changed
Introduces a novel space-time integrated architecture for robotics
Why It Matters
This model could significantly improve how embodied AI agents perceive and interact with physical environments by unifying spatial and temporal data processing.
What To Do Next
Monitor the project's GitHub or research papers for the release of their space-time architecture benchmarks.
Key Points
- •Introduces a novel space-time integrated architecture for robotics
- •Developed by a team with Fudan University academic background
- •Demonstrates strong market confidence with 5 funding rounds in 6 months
🧠 Deep Insight
Web-grounded analysis with 8 cited sources.
🔑 Enhanced Key Takeaways
- •The Fudan-linked team is identified as Moudian Intelligent, a company established in 2025.
- •Moudian Intelligent has successfully completed a 300 million RMB Pre-A funding round, marking their fourth funding round within six months, reflecting strong market confidence in foundational embodied intelligence technologies.
- •The company positions itself as the world's first edge-generative general embodied brain company, achieving full-chain connectivity across action generation, world modeling, and edge hardware-software collaboration.
- •A significant commercial milestone was reached in February 2026 with an order exceeding 100 million yuan to supply general brain modules for 10,000 humanoid robots intended for comprehensive life services.
- •Their innovative approach includes the 'World Motion Model' concept and a 'human-centered' embodied brain paradigm, featuring technologies like MotionGPT and HL3DWM (Human-Like 3D World Model).
📊 Competitor Analysis▸ Show
While direct comparative benchmarks for the Fudan-linked team's specific model are not available, the broader field of World Action Models (WAMs) and robot world models includes several key players:
| Feature/Company | Moudian Intelligent (Fudan-linked) | NVIDIA Cosmos WFMs | World Labs (Fei-Fei Li's) Marble | Tesla FSD | 1X World Model (1XWM) |
|---|---|---|---|---|---|
| Core Focus | Robot-native world action model, space-time integrated architecture, human-centered embodied brain paradigm, edge-generative general embodied brain. | Generative World Foundation Models (WFMs) for physical AI development, physics-aware synthetic video data generation. | Spatial intelligence, 3D comprehension, transforming single images into explorable worlds. | Largest deployed world model for autonomous driving, real-world validation at scale. | Video-pretrained world model, derives robot actions from text-conditioned video generation. |
| Architecture Highlights | HL3DWM (Human-Like 3D World Model) mimicking human 3D understanding, Object Perception Image Retrieval, Environmental Perception Information Aggregation. | Trained on 20M hours of robotics/driving data, open models with proprietary infrastructure. | Focus on 3D comprehension, building the “ImageNet of 3D worlds.” | Large-scale neural networks processing vast driving data, making life-or-death decisions. | World Model backbone (text-conditioned diffusion model) + Inverse Dynamics Model (IDM) for action sequence prediction. |
| Data Strategy | Leverages 3D point clouds and image details for global spatial relations. | 20 million hours of robotics and driving data. | Focus on 3D comprehension. | 3 billion+ miles of driving data from millions of vehicles. | Web-scale video, egocentric human data, NEO-specific sensorimotor logs. |
| Key Capabilities | Reusable abilities across bodies, scenes, and tasks; MotionGPT. | Generate photorealistic, physics-aware synthetic video data, simulate future scenarios. | Transform single images into explorable 3D worlds. | Real-time decision-making in complex driving scenarios. | Zero-shot generalization to novel objects, motions, and tasks without large-scale robot data pre-training. |
| Funding/Valuation | 300 million RMB Pre-A round, multiple rounds in 6 months. | Publicly traded company, provides infrastructure to competitors. | $1.25B valuation in 4 months (as of Aug 2025). | Publicly traded company, customers pay for FSD beta. | Backed by 1X Technologies. |
Other notable world model and VLA (Vision-Language-Action) model developers include Google's DeepMind (Genie 3, SIMA, Nano Banana), AMI Labs (Yann LeCun's VL-JEPA), Runway (GWM-1), Reactor, and Microsoft (Rho-alpha).
🛠️ Technical Deep Dive
- The Fudan-linked team's model, specifically the HL3DWM (Human-Like 3D World Model), is designed to mimic how humans understand the 3D world.
- This involves a process of first identifying relevant areas, then integrating surrounding information, and finally completing tasks.
- The architecture incorporates self-developed modules: an 'Object Perception Image Retrieval' module and an 'Environmental Perception Information Aggregation' module.
- These modules combine global spatial relations derived from 3D point clouds with fine details from images.
- This integration allows large language models to generate accurate answers and task solutions, even for complex tasks.
- Generally, World Action Models (WAMs) unify predictive state modeling with action generation, aiming for a joint distribution over future states and actions rather than just actions.
- WAMs often utilize large-scale video pretraining, including internet videos and egocentric human footage, to learn physical dynamics, which facilitates zero-shot generalization and human-to-robot transfer.
- For instance, NVIDIA's WAMs are described as a Joint Video-Action Diffusion Transformer (DiT) that concurrently predicts future latent visual tokens and corresponding robot actions, ensuring deep integration and reducing physically implausible actions.
- The 1X World Model (1XWM) employs a World Model backbone (a text-conditioned diffusion model trained on web-scale video, mid-trained on egocentric human data, and fine-tuned on robot sensorimotor logs) and an Inverse Dynamics Model (IDM) to extract necessary action sequences from generated frames.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗
