Xiaomi Streams Agent RL Training Live

Xiaomi is exposing the steps, rewards, tokens, and costs behind large-scale agent reinforcement learning.
30-Second TL;DR
What Changed
The stream covers MiMo-V2.6 Pro and MiMo-V2.6 Flash.
Why It Matters
Public training telemetry could improve reproducibility and help practitioners understand the economics of post-training agents. It may also encourage other labs to disclose more operational metrics.
What To Do Next
Monitor the MiMo RL dashboard and record reward-per-token and cost trends to compare against your own agent post-training runs.
Key Points
- •The stream covers MiMo-V2.6 Pro and MiMo-V2.6 Flash.
- •Viewers can inspect steps, token usage, rewards, and cost.
- •The dashboard offers uncommon visibility into agent RL at scale.
Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
Enhanced Key Takeaways
- •The training pipeline operates at a massive rollout scale, ingesting roughly 2 billion tokens across 1,568 prompts per step with 16 asynchronous rollouts per prompt.
- •Real-time financial tracking on the dashboard revealed cumulative spending surpassing $1 million to $1.13 million within 36-48 hours, with peak Pro runs burning approximately $20,000 per hour ($5 per second).
- •Live mid-training evaluations demonstrated that MiMo-V2.6-Pro reached 62% to nearly 66% on the DeepSWE v1.1 benchmark (mini-swe-agent, avg@3), up significantly from MiMo-V2.5's ~19% baseline.
- •The initiative is led by Xiaomi MiMo lead Luo Fuli (formerly of DeepSeek), who confirmed plans to progressively open-source the technical recipes, details, and infrastructure tools.
- •The agent RL effort runs strictly parallel to, but separate from, Xiaomi's concurrent Robotics-U0 embodied AI project (an autoregressive 38B/4B world model).
Competitor Analysis
- Xiaomi MiMo-V2.6 (Live RL Stream)
- Fully public live dashboard streaming metrics, spend, and restart logs in real time
- DeepSeek-R1
- Post-training recipes and weights released post-hoc via technical reports
- MiniMax-M1
- Closed training metrics; post-release reports/weights
- Xiaomi MiMo-V2.6 (Live RL Stream)
- Fully asynchronous, decoupled rollout/grader pipeline across 5 agent environments
- DeepSeek-R1
- Asynchronous multi-stage RL framework
- MiniMax-M1
- Distributed RL fine-tuning setup
- Xiaomi MiMo-V2.6 (Live RL Stream)
- 62% - ~66% (mid-training score)
- DeepSeek-R1
- Not directly evaluated on DeepSWE v1.1
- MiniMax-M1
- Not directly evaluated on DeepSWE v1.1
- Xiaomi MiMo-V2.6 (Live RL Stream)
- Surpassed $1M - $1.13M within 48h (~$20k/hr peak Pro burn)
- DeepSeek-R1
- Reported low single-digit million budget (not streamed live)
- MiniMax-M1
- Proprietary run budget (not streamed live)
| Feature / Metric | Xiaomi MiMo-V2.6 (Live RL Stream) | DeepSeek-R1 | MiniMax-M1 |
|---|---|---|---|
| Observability & Transparency | Fully public live dashboard streaming metrics, spend, and restart logs in real time | Post-training recipes and weights released post-hoc via technical reports | Closed training metrics; post-release reports/weights |
| RL Infrastructure Pipeline | Fully asynchronous, decoupled rollout/grader pipeline across 5 agent environments | Asynchronous multi-stage RL framework | Distributed RL fine-tuning setup |
| DeepSWE v1.1 Benchmark | 62% - ~66% (mid-training score) | Not directly evaluated on DeepSWE v1.1 | Not directly evaluated on DeepSWE v1.1 |
| Observed Post-Training Spend | Surpassed $1M - $1.13M within 48h (~$20k/hr peak Pro burn) | Reported low single-digit million budget (not streamed live) | Proprietary run budget (not streamed live) |
Technical Deep Dive
- Decoupled Asynchronous Rollout Pipeline: Operates fully asynchronously, decoupling rollout generation, environment execution, reward scoring, and parameter updates from rigid synchronization barriers.
- Scale & Throughput: Ingests ~2 billion tokens per training step across 1,568 prompts, generating 16 asynchronous rollouts per prompt.
- Multi-Task Agent Environments: Concurrently mixes five disparate agent environments—code, general reasoning, vision, chat, and cybersecurity—within a single unified training job.
- Novel Credit Assignment & Grader Compute: Incorporates agentic in-group credit assignment driven by deterministic unit test cases paired with rubric-based grader scoring.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-03Xiaomi MiMo team begins quiet experimentation on agent RL post-MiMo-V2.5
- 2026-09Xiaomi launches public streaming dashboard for MiMo-V2.6 Pro and Flash RL training
Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.


