arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASR 集成用于语音匿名化器的音素可懂度评估

ASR ensembling for phoneme intelligibility evaluation of speech anonymizers

Victor Ménestrel, Sebastian Möller, Slim Ouni, Dorothea Kolossa

arXiv 2609.28577首次发表:更新:

发表机构

Technische Universität Berlin; Université de Lorraine; CNRS; Inria; Loria(柏林工业大学; 洛林大学; 法国国家科学研究中心; 法国国家信息与自动化研究所; 洛林计算机科学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文首次对语音匿名化器进行音素级可懂度评估,提出基于ASR集成的硬投票指标,在组合多个模型时与人类评分相关性超0.9,并开源所有资源。

AI 中文摘要

我们首次对语音匿名化器进行了音素级可懂度评估,评估了基于ASR集成(自动语音识别集成)的指标与从众包听力测试中测得的可懂度之间的性能。我们的结果表明,只要组合多个ASR模型,简单的硬投票ASR指标在按特征、测试类型或条件聚合时,与人类评分的相关性可达到0.9以上;在有无载体句的情况下评估刺激,可进一步提高刺激级别的相关性。然而,基于后验概率的置信度指标并未带来增益,这可以追溯到此处使用的先进开源ASR模型校准不足。所有数据、代码和评估工具均已开源发布。

英文摘要

We present the first phoneme-level intelligibility evaluation of speech anonymizers, assessing the performance of ASR-ensemble-based metrics against measured intelligibility from a crowdsourced listening test. Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level. However, posterior-probability-based confidence metrics bring no gain, which can be traced back to the insufficient calibration of the state-of-the-art open ASR models that were utilized here. All data, code, and evaluation tools are released as open source.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑