arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

韵律风格能否仅从文本中推断?来自无监督声学聚类的证据

Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

Abdul Rehman, Jian-Jun Zhang, Xiaosong Yang

arXiv 2610.05575首次发表:更新:

发表机构

Bournemouth University(伯恩茅斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过声学聚类与文本嵌入预测实验,检验文本能否独立推断韵律风格,发现文本主要依赖词汇选择提供有限信息,且效果低于假设预期。

AI 中文摘要

大量表现性文本转语音研究基于一个未经检验的假设,即书面文本携带足够的信息来为其朗读选择适当的韵律风格。文本预测的风格模型改善了听者偏好,且表达恰当性评估预设了上下文约束风格,但两者均未直接衡量该假设本身。本文将这一假设作为可证伪的假设,针对仅从声学特征推导出的风格标签进行检验。在一个1200小时的对话语料库中,对六位说话者中的每一位,其话语在五个语音模型(包括一个仅韵律的对照组)的空间中进行聚类,并使用十二个文本嵌入模型预测留出话语的聚类。应用了三个对照条件:从语音嵌入中去除话语长度;准确率针对不平衡聚类的多数类基线(而非均匀随机概率)进行评分;以及一个词袋基线仅衡量词身份。文本预测聚类在所有六位说话者上均高于该基线(top-3准确率+0.111),但词袋方法达到了此效果的3/4。句子嵌入仅额外增加+0.026,其中最大的增益来自未针对句子语义训练的编码器,而基于树的探针在所有其他编码器上逆转了此增益。在360种配置中,声学聚类在文本嵌入空间中均不紧凑。仅韵律空间削弱了五位说话者的关联,但对表现最强的说话者并非如此。因此,文本主要通过词汇选择来提供这些朗读聚类的信息,无论是作为韵律的线索还是作为主题和录音情境的标志,且无参考的风格选择不能假设更多。

英文摘要

Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.

Comments13 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑