发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现最佳N选语音合成评估受ASR家族对齐影响,同家族验证器-评估器对效果更好。提出两种跨家族排名集成方法,能降低平均词错误率,建议交叉评估器三角测量作为默认报告实践。
AI 中文摘要
最佳N选(BoN)推理通过自动语音识别(ASR)验证器从N个候选中选择,提高了零样本文本到语音的内容一致性。我们发现一个未被充分探索的评估混淆因素:验证器的表观质量很大程度上取决于哪个ASR家族对其进行评判。在LibriSpeech-PC测试-clean上,使用F5-TTS时,验证器排名在Whisper、wav2vec 2.0和HuBERT评估器之间会发生反转。同家族验证器-评估器对比跨家族对能多恢复2-3倍的理想余量。我们提出了两种跨家族排名集成方法,在三个独立评估器中实现了最低平均词错误率,且在自动SIM-o/UTMOS指标下无明显退化。我们建议交叉评估器三角测量作为默认报告实践。
英文摘要
Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automatic speech recognition (ASR) verifier. We identify an evaluation confound: the apparent quality of a verifier depends strongly on the ASR family used for evaluation. On LibriSpeech-PC with F5-TTS, verifier rankings vary substantially across Whisper, wav2vec 2.0, and HuBERT evaluators, while same-family verifier and evaluator pairs recover considerably more oracle headroom than cross-family pairs despite highly similar representations. This pattern suggests identity- or lineage-level coupling rather than general representational similarity. To mitigate this bias, we propose two cross-family rank ensembles: rank averaging and conjunctive max-rank. Both improve mean word error rate across independent evaluators without degrading automatic similarity or quality metrics, and the best ensemble achieves a $12\%$ relative WER reduction over F5-TTS at $N=10$. These findings motivate cross-evaluator triangulation as a more reliable default for reporting BoN TTS performance.
CommentsAccepted at ICML 2026 Workshop on Machine Learning for Audio