发表机构
TTI-Chicago; Stony Brook University(芝加哥丰田技术研究所; 石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究量化了语音-文本语言模型与纯语音模型在生成上的模态差距,发现联合建模显著提升语义连贯性,且文本提供高效语义训练信号。
AI 中文摘要
纯语音语言模型在生成连贯内容方面往往落后于文本和语音-文本语言模型,但由于语音和文本系统通常使用不同的指标进行评估并在不同的数据上训练,这一差距难以量化。我们研究了一组基于流匹配进行连续声学特征生成的语音语言模型中的语音-文本模态差距。我们构建了一个统一的基于生成的评估套件,比较在匹配的数据分布上训练并在匹配的生成设置中评估的纯语音、纯文本和语音-文本语言模型。我们沿多个维度评估生成的续写:语义连贯性,通过转录生成的语音并使用参考语言模型进行评分来衡量;局部语音结构,通过音素n-gram分布统计来衡量;说话人一致性和声学质量;以及基于情感的分布指标。跨数据集,我们发现联合语音-文本建模显著提高了语义连贯性。然而,改进并非在所有指标上均匀分布:音素级指标变化不大,语音-文本续写的说话人相似性和预测质量较低,而基于情感的分布指标有所改善。与更大规模的纯语音模型相比,我们的语音-文本模型在基于转录的语义连贯性方面缩小了大部分扩展差距,这表明文本为语音语言建模提供了高效的语义训练信号。
英文摘要
Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and scoring it with a reference language model; local phonetic structure, measured by phone n-gram distributional statistics; speaker consistency and acoustic quality; and emotion-based distributional metrics. Across datasets, we find that joint speech-text modeling substantially improves semantic coherence. However, the improvement is not uniform across metrics: phone-level metrics change only modestly, speaker similarity and predicted quality are lower for speech-text continuations, while emotion-based distributional metrics improve. Compared with larger-scale speech-only models, our speech-text model closes much of the scaling gap in transcript-based semantic coherence, suggesting that text provides an efficient semantic training signal for spoken language modeling.
CommentsAccepted to SLT 2026, extended version with appendix