arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

响亮而清晰:动态激活引导提升嘈杂环境下的语音可懂度

Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

Seymanur Akti, Alexander Waibel

arXiv 2610.07647首次发表:更新:

发表机构

Karlsruhe Institute of Technology (KIT); KIT Campus Transfer (KCT); Carnegie Mellon University (CMU)(卡尔斯鲁厄理工学院; KIT 校园转移中心; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种无需重训练的激活引导方法,通过动态调整发声努力和过度发音,提升TTS在嘈杂环境下的语音可懂度,保持说话人相似度并显著降低词错误率。

AI 中文摘要

在嘈杂环境中,语音的可懂度会降低,而人类会自然地调整自己的声音来进行补偿。受此行为启发,我们研究能否通过激活引导(activation steering)来引导文本到语音(TTS)模型生成更可懂的语音,而无需重新训练。我们聚焦于Lombard效应的两个特征:增强的发声努力和过度发音(hyper-articulation)。我们提出了一种提示相对引导(prompt-relative steering)机制,该机制可防止引导效应在生成过程中累积,同时允许动态调整其强度。在见过的和未见过的说话人以及多种语言上,我们的方法在Lombard相关声学特征上产生了系统性变化,保持了说话人相似度(89-95%),并在1 dB信噪比(SNR)下将背景噪声下的词错误率(WER)降低了7-22%。这些结果表明,预训练的TTS模型可以被动态控制以生成更可懂的语音,而无需重新训练。

英文摘要

Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑