SourceRecentcollected in 15h

Xiaomi Streams Agent RL Training Live

Read original on Pandaily
#agent-training#training-costs

Xiaomi is exposing the steps, rewards, tokens, and costs behind large-scale agent reinforcement learning.

30-Second TL;DR

What Changed

The stream covers MiMo-V2.6 Pro and MiMo-V2.6 Flash.

Why It Matters

Public training telemetry could improve reproducibility and help practitioners understand the economics of post-training agents. It may also encourage other labs to disclose more operational metrics.

What To Do Next

Monitor the MiMo RL dashboard and record reward-per-token and cost trends to compare against your own agent post-training runs.

Who should care:Researchers & Academics

Key Points

  • The stream covers MiMo-V2.6 Pro and MiMo-V2.6 Flash.
  • Viewers can inspect steps, token usage, rewards, and cost.
  • The dashboard offers uncommon visibility into agent RL at scale.
Key numbers$1 million$1.13 million$20,000$5

Deep Insight

Background and context from public sources — not the original article. 7 sources cited.

Enhanced Key Takeaways

  • The training pipeline operates at a massive rollout scale, ingesting roughly 2 billion tokens across 1,568 prompts per step with 16 asynchronous rollouts per prompt.
  • Real-time financial tracking on the dashboard revealed cumulative spending surpassing $1 million to $1.13 million within 36-48 hours, with peak Pro runs burning approximately $20,000 per hour ($5 per second).
  • Live mid-training evaluations demonstrated that MiMo-V2.6-Pro reached 62% to nearly 66% on the DeepSWE v1.1 benchmark (mini-swe-agent, avg@3), up significantly from MiMo-V2.5's ~19% baseline.
  • The initiative is led by Xiaomi MiMo lead Luo Fuli (formerly of DeepSeek), who confirmed plans to progressively open-source the technical recipes, details, and infrastructure tools.
  • The agent RL effort runs strictly parallel to, but separate from, Xiaomi's concurrent Robotics-U0 embodied AI project (an autoregressive 38B/4B world model).

Competitor Analysis

Observability & Transparency
Xiaomi MiMo-V2.6 (Live RL Stream)
Fully public live dashboard streaming metrics, spend, and restart logs in real time
DeepSeek-R1
Post-training recipes and weights released post-hoc via technical reports
MiniMax-M1
Closed training metrics; post-release reports/weights
RL Infrastructure Pipeline
Xiaomi MiMo-V2.6 (Live RL Stream)
Fully asynchronous, decoupled rollout/grader pipeline across 5 agent environments
DeepSeek-R1
Asynchronous multi-stage RL framework
MiniMax-M1
Distributed RL fine-tuning setup
DeepSWE v1.1 Benchmark
Xiaomi MiMo-V2.6 (Live RL Stream)
62% - ~66% (mid-training score)
DeepSeek-R1
Not directly evaluated on DeepSWE v1.1
MiniMax-M1
Not directly evaluated on DeepSWE v1.1
Observed Post-Training Spend
Xiaomi MiMo-V2.6 (Live RL Stream)
Surpassed $1M - $1.13M within 48h (~$20k/hr peak Pro burn)
DeepSeek-R1
Reported low single-digit million budget (not streamed live)
MiniMax-M1
Proprietary run budget (not streamed live)

Technical Deep Dive

  • Decoupled Asynchronous Rollout Pipeline: Operates fully asynchronously, decoupling rollout generation, environment execution, reward scoring, and parameter updates from rigid synchronization barriers.
  • Scale & Throughput: Ingests ~2 billion tokens per training step across 1,568 prompts, generating 16 asynchronous rollouts per prompt.
  • Multi-Task Agent Environments: Concurrently mixes five disparate agent environments—code, general reasoning, vision, chat, and cybersecurity—within a single unified training job.
  • Novel Credit Assignment & Grader Compute: Incorporates agentic in-group credit assignment driven by deterministic unit test cases paired with rubric-based grader scoring.

Future ImplicationsAI analysis grounded in cited sources

Frontier AI labs will face heightened pressure to open-source post-training metrics and live run economics.
Xiaomi's live exposure of failure logs, rollout counts, and financial burn sets an unprecedented public transparency standard that challenges the opaque, closed-door post-training practices of frontier competitors.
Xiaomi will emerge as a top-tier open-weight reasoning and agentic model provider.
The demonstrated leap on the DeepSWE v1.1 benchmark from ~19% to over 62% signals that MiMo-V2.6's upcoming open-source recipe could rapidly disseminate frontier agentic coding capabilities to the broader developer community.

Timeline

2026-03
Xiaomi MiMo team begins quiet experimentation on agent RL post-MiMo-V2.5
2026-09
Xiaomi launches public streaming dashboard for MiMo-V2.6 Pro and Flash RL training

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.