发表机构
Harbin Institute of Technology; Zhejiang University(哈尔滨工业大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对基于语言模型的文本转语音的语音幻觉问题,提出无训练的经验校准对比解码方法,在多数据集上显著降低错误率并提升听辨质量,为解码时缓解语音幻觉提供新方向。
AI 中文摘要
基于语言模型的文本转语音(LM-based TTS)仍易出现偏离目标文本的语音幻觉。现有缓解方法主要依赖架构变更或额外训练,而解码时的控制方法研究不足。本文提出一种条件信息视图,将文本衍生的对齐信息与声学上下文及学习到的语音规律提供的经验信息区分开。我们假设,当对齐支持在脆弱转换点未充分反映在所选 token 中时,会出现一类重要的幻觉。利用同一语音 LM 在有文本条件和无文本条件下的预测结果,我们提出经验校准对比解码(Experience-Calibrated Contrastive Decoding, ECCD),这是一种无需训练的方法,可在保留有用经验信息的同时增强对齐支持。ECCD 保留原始专家分布,仅应用正向对齐增强,并利用集合级经验兼容性校准其强度。在四个模型上,ECCD 在所有 SeedTTS-Eval 设置中使词错误率(WER)/字符错误率(CER)降低高达 55.6%,在 25 种多语言 CV3-Eval 设置中的 24 种设置中实现降低。听辨测试获得平均意见分差(CMOS)提升 +0.644,同时保持强说话人相似度。进一步分析显示,对齐影响和决策级增益在语言单元内存在差异,且在首次错误边界处的影响低于匹配正确边界处。总体而言,这些广泛的实验和分析表明,条件信息控制是缓解语音幻觉的一个有前景的解码时研究方向。
英文摘要
Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text. Existing mitigation mainly relies on architectural changes or additional training, while decoding-time control remains underexplored. We present a conditional information view that distinguishes text-derived alignment information from experience information supplied by acoustic context and learned speech regularities. We hypothesize that an important class of hallucinations begins when alignment support is insufficiently reflected in the selected token at a vulnerable transition. Using predictions from the same speech LM with and without text conditions, we propose Experience-Calibrated Contrastive Decoding (ECCD), a training-free method that strengthens alignment support while preserving useful experience information. ECCD preserves the original expert distribution, applies only positive alignment enhancement, and calibrates its strength using set-level experience compatibility. Across four models, ECCD reduces WER/CER by up to 55.6% in all SeedTTS-Eval settings and 24 of 25 multilingual CV3-Eval settings. A listening test yields a CMOS gain of $+0.644$ while retaining strong speaker similarity. Further analysis shows that alignment influence and decision-level gain vary within linguistic units and are lower at first-error boundaries than at matched correct boundaries. Overall, these extensive experiments and analyses identify conditional information control as a promising decoding-time direction for mitigating speech hallucination.
CommentsWork in progress