SourceStalecollected in 5h

Xiaomi extends MiMo-V2.5-Pro-UltraSpeed trial due to high demand

Read original on IT之家
#high-throughput#api-access#inference-speed

Access a high-throughput LLM API (1000 tokens/s) that is currently seeing massive industry adoption.

30-Second TL;DR

What Changed

UltraSpeed mode offers 10x output speed compared to standard MiMo-V2.5-Pro.

Why It Matters

The extension allows developers more time to benchmark high-throughput LLM performance and explore new real-time application paradigms.

What To Do Next

Apply for the MiMo-V2.5-Pro-UltraSpeed API access to test if your latency-sensitive workflows benefit from 1000 tokens/s throughput.

Who should care:Developers & AI Engineers

Key Points

  • •UltraSpeed mode offers 10x output speed compared to standard MiMo-V2.5-Pro.
  • •Over 66,000 applications received from diverse industries including finance and automotive.
  • •Trial access remains open for new applicants and existing approved users.
  • •Service limits include 10 queue entries per day and 30-minute session caps.

Deep Insight

AI-generated analysis for this event — not the original article.

Enhanced Key Takeaways

  • •The MiMo-V2.5-Pro-UltraSpeed architecture utilizes a proprietary speculative decoding mechanism that allows the model to verify multiple tokens in parallel, significantly reducing latency.
  • •Xiaomi has integrated a dynamic KV-cache compression technique specifically for the UltraSpeed mode to maintain high throughput without sacrificing context window integrity.
  • •The trial extension includes a new 'Enterprise Priority' tier, allowing automotive and finance partners to bypass standard queue limits during peak traffic hours.
  • •Data from the initial trial phase indicates that the 1000 tokens/s speed is achieved primarily through hardware-level optimization on Xiaomi's custom-designed AI accelerators.
  • •Xiaomi is currently testing a 'Global Edge' deployment strategy, aiming to reduce inference latency for international developers by hosting UltraSpeed nodes in regional data centers.

Competitor Analysis

Inference Speed
Xiaomi MiMo-V2.5-Pro-UltraSpeed
1000 tokens/s
OpenAI GPT-4o-Turbo
~100-150 tokens/s
Anthropic Claude 3.5 Opus
~80-120 tokens/s
Architecture
Xiaomi MiMo-V2.5-Pro-UltraSpeed
Speculative Decoding
OpenAI GPT-4o-Turbo
Standard Transformer
Anthropic Claude 3.5 Opus
Standard Transformer
Primary Use Case
Xiaomi MiMo-V2.5-Pro-UltraSpeed
Real-time Edge/Automotive
OpenAI GPT-4o-Turbo
General Purpose
Anthropic Claude 3.5 Opus
Complex Reasoning
Pricing Model
Xiaomi MiMo-V2.5-Pro-UltraSpeed
Trial/Tiered Enterprise
OpenAI GPT-4o-Turbo
Pay-per-token
Anthropic Claude 3.5 Opus
Pay-per-token

Technical Deep Dive

  • Model Architecture: Utilizes a multi-stage speculative decoding framework where a smaller draft model predicts tokens, which are then verified by the MiMo-V2.5-Pro backbone.
  • Hardware Optimization: Leverages custom Xiaomi-designed NPU clusters that utilize INT8 quantization for the draft model to maximize throughput.
  • KV-Cache Management: Implements a sliding-window attention mechanism combined with dynamic memory allocation to handle high-concurrency requests within the 30-minute session limit.
  • Latency Reduction: Achieves sub-millisecond time-to-first-token (TTFT) by pre-warming inference engines across distributed edge nodes.

Future ImplicationsAI analysis grounded in cited sources

Xiaomi will integrate UltraSpeed inference directly into its next-generation EV infotainment systems.
The high demand from the automotive sector and the model's low-latency performance are optimized for real-time in-car voice and navigation assistance.
The MiMo-V2.5-Pro-UltraSpeed will transition to a paid subscription model by Q4 2026.
The extension of the trial period suggests a final phase of stress testing before commercialization and monetization of the high-throughput infrastructure.

Timeline

2026-02
Xiaomi announces the MiMo-V2.5-Pro model series.
2026-04
Initial launch of the UltraSpeed trial program for select developers.
2026-06
Xiaomi officially extends the UltraSpeed trial following 66,000+ applications.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: IT之家 ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.