口语语言模型在阅读文本时是否“听到”语音?弥合语音与文本之间的结构差距
Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text
浏览论文内容
中文总结 AI 辅助
该研究针对现有口语语言模型(SLMs)未充分解决语音与文本结构差异的问题,提出解耦长度不匹配与语义对齐的框架,经多基准实验验证其性能可与强基线媲美,凸显了明确解决语音文本结构差异的重要性。
中文摘要 AI 辅助
口语语言模型(SLMs)可直接从语音生成文本响应,为级联系统提供了替代方案。尽管近期取得了进展,现有SLMs与基于文本的语言模型相比,在指令遵循行为和跨不同任务的泛化能力上仍较弱。我们的分析表明,尽管当前SLMs的下游性能较强,但其中的语音和文本表示仍对齐度较低,这表明连续、随时间变化的语音与离散文本之间的结构差异仍未得到充分解决。为解决该问题,我们提出了一个简单框架,将长度不匹配与语义解耦,并鼓励语音和文本表示之间更紧密的对应关系。在多个基准上开展的实验显示,该框架的性能与强基线具有竞争力,凸显了在SLM训练中明确解决语音与文本之间结构差异的重要性。我们的代码可在this https URL获取。
英文摘要
Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited generalization across diverse tasks compared to text-based language models. Our analysis shows that speech and text representations in current SLMs remain weakly aligned despite strong downstream performance, indicating that structural differences between continuous, temporally varying speech and discrete text remain insufficiently addressed. To address this, we propose a simple framework that decouples length mismatch from semantic alignment and encourages closer correspondence between speech and text representations. Experiments across multiple benchmarks demonstrate competitive performance against strong baselines, underscoring the importance of explicitly addressing structural differences between speech and text in SLM training. Our code is publicly available at https://github.com/jaykim9870/Do_SLMs_Hear_Speech_as_They_Read_Text.
发表机构
- Maum AI Inc.(Maum AI公司)
- KAIST(韩国科学技术院)
- Atmanity Inc.(Atmanity公司)
机构由 AI 辅助整理,请以论文原文为准。