來源虎嗅•較早收集於 19m
MOSS團隊轉型:專注端到端多模態語音AI

了解 MOSS 團隊為何放棄純文本模型,轉向端到端語音以解決實時交互延遲問題。
30 秒速覽
有什麼變化
從純文本大模型轉向端到端語音優先的多模態模型,以降低延遲並減少信息丟失。
為什麼重要
通過專注於端到端語音處理,該團隊旨在解決傳統級聯式(ASR-LLM-TTS)架構中固有的延遲與情感細節缺失問題,可能為 AI 硬體集成樹立新的標準。
下一步行動
在為硬體構建實時交互式 AI 代理時,評估端到端語音模型與級聯式架構在性能上的權衡。
誰應關注:Developers & AI Engineers
關鍵要點
- •從純文本大模型轉向端到端語音優先的多模態模型,以降低延遲並減少信息丟失。
- •開發「情境智能」,使模型能感知物理環境、情緒及說話人身份。
- •產品線涵蓋 MOSS-Transcribe-Diarize、MOSS-Video-Preview 及 MOSS-TTS。
- •採取用戶導向的開發模式,根據實際場景需求(如時延、性能)反推模型架構與訓練方案。
深度解析
本篇為 AI 生成分析,非原文內容。
增強重點摘要
- •Moss Intelligence was established by Professor Qiu Xipeng and his team from Fudan University, transitioning the academic MOSS project into a commercial entity.
- •The company has secured strategic backing from major Chinese venture capital firms focusing on AGI infrastructure to support the high compute costs of end-to-end multimodal training.
- •The 'Situational Intelligence' framework utilizes a proprietary streaming architecture that processes audio tokens directly without intermediate text conversion, significantly reducing 'time-to-first-token' latency.
- •Moss Intelligence is targeting the enterprise 'Digital Employee' market, specifically focusing on high-stakes sectors like legal transcription and real-time medical diagnostics where speaker diarization accuracy is critical.
- •The team is leveraging a hybrid training strategy that combines large-scale synthetic audio data with real-world acoustic environment datasets to improve robustness in noisy, non-studio conditions.
競品分析
Architecture
- Moss Intelligence
- End-to-End Native
- OpenAI (Voice Mode)
- End-to-End Native
- DeepSeek (Voice)
- Text-to-Speech Pipeline
Primary Focus
- Moss Intelligence
- Situational/Physical Context
- OpenAI (Voice Mode)
- Conversational Fluidity
- DeepSeek (Voice)
- Reasoning/Coding
Diarization
- Moss Intelligence
- High-Precision Native
- OpenAI (Voice Mode)
- Standard
- DeepSeek (Voice)
- Limited
Market Target
- Moss Intelligence
- Enterprise/Industrial
- OpenAI (Voice Mode)
- Consumer/General
- DeepSeek (Voice)
- Developer/API
| Feature | Moss Intelligence | OpenAI (Voice Mode) | DeepSeek (Voice) |
|---|---|---|---|
| Architecture | End-to-End Native | End-to-End Native | Text-to-Speech Pipeline |
| Primary Focus | Situational/Physical Context | Conversational Fluidity | Reasoning/Coding |
| Diarization | High-Precision Native | Standard | Limited |
| Market Target | Enterprise/Industrial | Consumer/General | Developer/API |
技術深入
- Architecture: Utilizes a unified transformer backbone that processes multi-stream inputs (audio, visual, text) into a shared latent space.
- Latency Optimization: Implements a speculative decoding mechanism specifically tuned for audio tokens to maintain sub-200ms response times.
- Diarization Engine: Employs a speaker-embedding module integrated directly into the attention layers, allowing the model to maintain speaker identity across long-form audio sessions.
- Training Data: Employs a multi-stage curriculum learning approach, starting with synthetic speech-to-intent tasks and scaling to complex, multi-speaker situational dialogues.
前景展望基於引用來源的 AI 分析
Moss Intelligence will achieve a dominant market share in the Chinese enterprise voice-AI sector by 2027.
Their focus on specialized, high-accuracy diarization and situational awareness addresses specific pain points in Chinese enterprise workflows that general-purpose models currently struggle with.
The shift to end-to-end voice models will trigger a consolidation of the Chinese TTS and transcription software market.
As end-to-end models provide superior latency and emotional nuance, legacy pipeline-based transcription and TTS providers will face significant obsolescence pressure.
時間線
2023-02
Fudan University releases MOSS, the first Chinese conversational LLM, to the public.
2023-04
MOSS project open-sources its model weights and codebase on GitHub.
2025-09
Moss Intelligence is officially incorporated to commercialize multimodal voice technologies.
2026-03
Company announces the successful training of its first end-to-end multimodal foundation model.
- 2023-02Fudan University releases MOSS, the first Chinese conversational LLM, to the public.
- 2023-04MOSS project open-sources its model weights and codebase on GitHub.
- 2025-09Moss Intelligence is officially incorporated to commercialize multimodal voice technologies.
- 2026-03Company announces the successful training of its first end-to-end multimodal foundation model.
AI 週報
閱讀本週精選 AI 大事摘要 →
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 虎嗅 ↗
每週電子報
每週一封,可隨時退訂。



