DM0.5 Tops RoboDojo with Full Open Source

💡See how an open-source embodied model reached the top of RoboDojo.
⚡ 30-Second TL;DR
What Changed
DM0.5 achieved the top position on the RoboDojo benchmark.
Why It Matters
A leading result on RoboDojo could raise expectations for embodied AI control and generalization. Full open sourcing may accelerate independent reproduction and competition in robot foundation models.
What To Do Next
Clone the released DM0.5 code and reproduce its RoboDojo evaluation before adapting it to your robot platform.
Key Points
- •DM0.5 achieved the top position on the RoboDojo benchmark.
- •The system is presented as a new state-of-the-art robot brain.
- •Its full open-source availability enables cloning, testing, and further research.
🧠 Deep Insight
Background and context from public sources — not the original article. 15 sources cited.
🔑 Enhanced Key Takeaways
- •DM0.5 is built upon a Gemma3-4B VLM base architecture, integrated with a specialized 680M parameter Flow-Matching action expert.
- •The model was trained on a massive dataset of 150,000 hours of robot interaction data, representing a 400% increase over the previous DM0 iteration.
- •The system supports advanced embodied reasoning capabilities, specifically featuring 60-second historical context abstraction and 11 distinct types of embodied Chain-of-Thought (CoT).
- •Dexmal (Yuanli Lingji) was founded in March 2025 by a core team originating from Megvii Technology, focusing on a Model-as-a-Service (MaaS) business model.
- •The RoboDojo benchmark, where DM0.5 achieved its top ranking, is a collaborative effort involving nearly 20 global institutions, including UC Berkeley and Tsinghua, covering 42 simulation and 18 real-world tasks.
📊 Competitor Analysis▸ Show
| Feature | DM0.5 (Dexmal) | RT-2 (Google DeepMind) | OpenVLA |
|---|---|---|---|
| Architecture | Gemma3-4B + Flow-Matching | PaLM-E / ViT-based | Llama-2-7B + DINOv2 |
| Open Source | Full (Weights/Scripts) | Proprietary | Full |
| Training Data | 150,000 hours | Proprietary Web/Robot | Open-X Embodiment |
| Primary Focus | Industrial/Precision | General Purpose | Research/Academic |
🛠️ Technical Deep Dive
- Model Architecture: 4B-parameter Vision-Language-Action (VLA) model.
- Base Model: Gemma3-4B VLM.
- Action Head: 680M parameter Flow-Matching action expert.
- Context Window: Supports 60-second historical context abstraction.
- Reasoning: Implements 11 types of embodied Chain-of-Thought (CoT) for task planning.
- Hardware Compatibility: Includes modification guides for AgileX COBOT Magic and other robotic platforms.
- Ecosystem: Provides fine-tuned checkpoints for LIBERO, RoboTwin2.0, and SO101 benchmarks.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.

