Search

Tag: #multimodal-llm14 results

Closing Text-Speech Gap in LLMs

Closing Text-Speech Gap in LLMs

Speech-adapted LLMs underperform text-based counterparts and cascaded pipelines on language tasks, termed the text-speech understanding gap. This performance drop occurs when processing spoken inputs versus equivalent text. Recent gap-narrowing methods rely on costly large-scale speech synthesis.

Apple Machine LearningOfficialFeb 25#speech-understanding#performance-gap#multimodal-llm
Page 2 of 2