发表机构
Technische Universität Berlin; Université de Lorraine; CNRS; Inria; Loria(柏林工业大学; 洛林大学; 法国国家科学研究中心; 法国国家信息与自动化研究所; 洛林计算机科学实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次对语音匿名化器进行音素级可懂度评估,提出基于ASR集成的硬投票指标,在组合多个模型时与人类评分相关性超0.9,并开源所有资源。
AI 中文摘要
我们首次对语音匿名化器进行了音素级可懂度评估,评估了基于ASR集成(自动语音识别集成)的指标与从众包听力测试中测得的可懂度之间的性能。我们的结果表明,只要组合多个ASR模型,简单的硬投票ASR指标在按特征、测试类型或条件聚合时,与人类评分的相关性可达到0.9以上;在有无载体句的情况下评估刺激,可进一步提高刺激级别的相关性。然而,基于后验概率的置信度指标并未带来增益,这可以追溯到此处使用的先进开源ASR模型校准不足。所有数据、代码和评估工具均已开源发布。
英文摘要
We present the first phoneme-level intelligibility evaluation of speech anonymizers, assessing the performance of ASR-ensemble-based metrics against measured intelligibility from a crowdsourced listening test. Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level. However, posterior-probability-based confidence metrics bring no gain, which can be traced back to the insufficient calibration of the state-of-the-art open ASR models that were utilized here. All data, code, and evaluation tools are released as open source.