🦙Stalecollected in 78m

Xiaomi MiMo V2 Pro Benchmarks Leaked

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#benchmark-leak#model-previewmimo-v2-proxiaomimimo-v2-pro

💡Early benchmarks of Xiaomi's new local LLM rival—check performance vs leaders

⚡ 30-Second TL;DR

What Changed

Leak originates from Reddit r/LocalLLaMA post

Why It Matters

This leak provides early insights into Xiaomi's AI model performance, potentially influencing competitor strategies and local LLM adoption.

What To Do Next

Visit artificialanalysis.ai/models/mimo-v2-pro to review leaked MiMo V2 Pro benchmarks.

Who should care:Researchers & Academics

Key Points

  • Leak originates from Reddit r/LocalLLaMA post
  • Benchmarks accessible at artificialanalysis.ai/models/mimo-v2-pro
  • Submitted by user /u/External_Mood4719
  • Focuses on Xiaomi's new MiMo V2 Pro model

🧠 Deep Insight

Background and context from public sources — not the original article. 6 sources cited.

🔑 Enhanced Key Takeaways

  • MiMo-V2-Flash, a related model in Xiaomi's lineup released in February 2026, is a Mixture-of-Experts (MoE) architecture with 309B total parameters and 15B active parameters[3].
  • MiMo-V2-Flash achieves top rankings among open-source models on AIME 2025 math competition and GPQA-Diamond scientific benchmark[3].
  • The model supports infinite context windows and delivers inference speeds up to 150 tokens per second at a cost of $0.1 per million input tokens[3][2].

🛠️ Technical Deep Dive

  • MiMo-V2-Flash uses a hybrid attention mechanism with a 1:5 ratio of Global Attention (GA) and Sliding Window Attention (SWA), featuring an aggressive 128-token sliding window for efficiency[3].
  • Includes a lightweight Multi-Token Prediction (MTP) block with a dense FFN and SWA, achieving 2.0–2.6× speedup and accepted lengths of 2.8–3.6 tokens[3].
  • Demonstrates near 100% long-context retrieval success from 32K to 256K tokens and robust performance on GSM-Infinite benchmark up to 128K[5].
  • Benchmark scores include 83.5% on GPQA, 33.5 coding index, and 67.7% on AIME 2025[2].

🔮 Future ImplicationsAI analysis grounded in cited sources

MiMo V2 Pro will prioritize efficiency in MoE scaling
Building on V2-Flash's 309B/15B MoE design and hybrid attention, leaked benchmarks suggest continued focus on active parameter optimization for speed and cost[3].
Xiaomi models will challenge top open-source leaders in reasoning
V2-Flash's top-2 open-source rankings on GPQA-Diamond and AIME indicate V2 Pro leaks aim to extend this edge in analytical tasks[3][2].

Timeline

2026-02
Xiaomi releases MiMo-V2-Flash with MoE architecture and high benchmark scores
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.