Mistral Launches Local Japanese Speech AI Voxtral 2

💡Mistral's local Japanese speech model: offline batch/real-time options beat cloud costs for devs.
⚡ 30-Second TL;DR
What Changed
Mistral AI released Voxtral Transcribe 2 speech recognition model
Why It Matters
This launch enables offline, privacy-preserving speech-to-text for Japanese applications, lowering costs for batch tasks and enabling responsive real-time use without cloud reliance. AI practitioners can integrate it into edge devices for cost-effective solutions.
What To Do Next
Download Voxtral Transcribe 2 from Mistral's Hugging Face repo and test local Japanese batch transcription.
Key Points
- •Mistral AI released Voxtral Transcribe 2 speech recognition model
- •Supports Japanese language and runs entirely locally
- •Two models: high-accuracy low-price batch processing variant
- •Ultra-low latency real-time transcription variant
🧠 Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
🔑 Enhanced Key Takeaways
- •Voxtral Transcribe 2 natively supports 13 languages including Chinese, Arabic, Russian, Korean, and Japanese alongside others.
- •Voxtral Mini Transcribe V2 achieves a 4% word error rate on the FLEURS benchmark and outperforms GPT-4o mini, Gemini 2.5 Flash, Assembly Universal, and Deepgram Nova in accuracy.
- •Pricing for Voxtral Transcribe 2 is $0.003 per minute, processing audio 3x faster than ElevenLabs Scribe v2 at one-fifth the cost.
- •Models support context biasing (optimized for English), low susceptibility to background noise, and speaker recognition for enterprise use cases like call centers.
📊 Competitor Analysis▸ Show
| Feature | Voxtral Transcribe 2 / Voxtral Mini V2 | GPT-4o mini | Gemini 2.5 Flash | ElevenLabs Scribe v2 | Deepgram Nova |
|---|---|---|---|---|---|
| WER on FLEURS | ~4% | Outperformed | Outperformed | Comparable quality | Outperformed |
| Languages | 13 (incl. Japanese, Chinese, Arabic) | - | - | - | - |
| Speed | 3x faster than Scribe v2 | - | - | Baseline | - |
| Pricing | $0.003/min | - | - | 5x higher | - |
| Speaker Recognition | Yes | - | - | - | - |
🛠️ Technical Deep Dive
- •Built on Mistral Small 3.1 LLM backbone; Voxtral Small variant has 24B parameters, smaller variant 3B for faster inference.
- •Supports up to 30min transcription or 40min audio understanding with 32k token context; automatic language detection.
- •Audio playground in Mistral Studio allows uploading up to 10 files (MP3, WAV, M4A, FLAC, OGG; max 1GB/file) with speaker recognition, timestamps, and context biasing.
- •Enterprise features: private deployment, domain-specific fine-tuning, function-calling from voice for API triggers.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- TechCrunch — Mistral Releases Voxtral Its First Open Source AI Audio Model
- benzatine.com — Mistral Unveils Voxtral a Game Changer in Open Source Voice AI
- GitHub — Cog Mistralai Voxtral
- trendingtopics.eu — Mistral AI Launches Voxtral Transcribe 2 for Real Time Speech Recognition
- mistral.ai — Voxtral
- gigazine.net — 20251219 Mistral Ocr 3
- mistral.ai — Mistral Ocr
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.


