⚛️量子位•Stalecollected in 61m
Jieyue Tops China Speech Benchmark

💡China's #1 speech model tops benchmark—key for audio AI devs!
⚡ 30-Second TL;DR
What Changed
Jieyue's new speech model leads Artificial Analysis China leaderboard
Why It Matters
Strengthens China's speech AI competitiveness, potentially accelerating local R&D and adoption in voice applications.
What To Do Next
Benchmark your speech models on Artificial Analysis leaderboard against Jieyue's top performer.
Who should care:Researchers & Academics
Key Points
- •Jieyue's new speech model leads Artificial Analysis China leaderboard
- •First place among all Chinese speech models evaluated
- •Benchmark focuses on key voice AI metrics
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •Jieyue (StepFun) achieved this ranking through its 'Step-2' multimodal model architecture, which integrates native speech-to-speech capabilities rather than relying on traditional cascaded ASR-LLM-TTS pipelines.
- •The Artificial Analysis benchmark specifically highlighted Jieyue's low latency and high emotional expressiveness, which outperformed domestic competitors like Qwen-Audio and Yi-Audio in real-time conversational scenarios.
- •This performance milestone marks a strategic shift for Jieyue, moving from general-purpose text-based LLMs to specialized, high-fidelity multimodal interaction models designed for edge-device integration.
📊 Competitor Analysis▸ Show
| Feature | Jieyue (Step-2) | Qwen-Audio (Alibaba) | Yi-Audio (01.AI) |
|---|---|---|---|
| Architecture | Native Multimodal | Cascaded/Hybrid | Cascaded |
| Latency | Ultra-low (End-to-End) | Moderate | Moderate |
| Primary Focus | Real-time Voice Interaction | General Audio Understanding | Speech Recognition/Synthesis |
| Benchmark Rank | #1 (China) | #3 (China) | #5 (China) |
🛠️ Technical Deep Dive
- •Architecture: Utilizes a unified transformer-based backbone that processes audio tokens directly alongside text tokens, eliminating the need for intermediate text transcription.
- •Inference Optimization: Employs proprietary quantization techniques to maintain high-fidelity audio output while reducing VRAM requirements for deployment on consumer-grade hardware.
- •Training Data: Leveraged a massive, proprietary dataset of high-quality, emotionally nuanced human-to-human conversational audio, specifically curated for Chinese linguistic nuances.
- •Latency Metrics: Achieved a Time-to-First-Token (TTFT) for audio generation under 200ms in controlled testing environments.
🔮 Future ImplicationsAI analysis grounded in cited sources
Jieyue will capture significant market share in the Chinese smart automotive and IoT voice assistant sectors by Q4 2026.
The model's native low-latency performance provides a distinct competitive advantage for real-time, in-cabin voice control systems where traditional cascaded models suffer from lag.
Domestic Chinese AI labs will shift development focus toward native multimodal speech architectures by early 2027.
Jieyue's benchmark success validates that end-to-end audio processing is the new performance ceiling, forcing competitors to abandon legacy ASR-LLM-TTS pipelines to remain relevant.
⏳ Timeline
2023-07
StepFun (Jieyue) founded by former Microsoft and Google AI researchers.
2024-03
Release of Step-1, the company's first general-purpose large language model.
2025-05
Introduction of Step-2, the company's flagship multimodal model with native audio capabilities.
2026-05
Jieyue's speech model secures top ranking on Artificial Analysis China leaderboard.
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗