发表机构
Institute of Software, Chinese Academy of Sciences; Capital Normal University(中国科学院软件研究所; 首都师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过两个实验比较自然与合成语音的语调感知,发现语音类型与语调及熟悉度存在交互作用,为语音克隆在语言训练语料构建中的应用提供初步评估。
AI 中文摘要
语言训练依赖于由大量语言材料构成的语料库。AI驱动的语音克隆提供了一种以相对较低成本构建语料库的方法。歌唱声音转换(SVC)模型用于生成合成语音。本研究通过两个实验(相似性感知和语调识别)比较了参与者在自然语音和合成语音上的表现。在相似性感知任务的准确率中,发现语音类型与语调之间存在显著交互作用,表明疑问句可能作为说话人识别的线索,但可能受到合成特征的影响。在语调识别任务的准确率中,观察到语音类型与熟悉度之间存在显著交互作用,表明语音类型影响熟悉度对语音处理的贡献程度。
英文摘要
Language training relies on a corpus constructed by a large number linguistic materials. AI-powered voice clones provide a way to construct the corpus with relatively low cost. Singing voice conversion (SVC) model is used to generate synthetic voices. This study compares participants' performances on natural and synthetic speech in two experiments, similarity perception and intonation recognition. In the accuracy of similarity perception task, a significant interaction between speech type and intonation is found, suggesting that question may serve as a cue for speaker identification but may be influenced by synthetic features. In the accuracy of intonation recognition task, a significant interaction between speech type and familiarity is observed, indicating that speech type affects how much familiarity contributes to voice processing.
CommentsAccepted by Interspeech 2026