arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进预测的MOS分数,而非感知质量:增强语音的多预测器测试时优化

Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech

Tsubasa Ochiai, Marc Delcroix, Nahomi Kusunoki, Rintaro Ikeshita, Naohiro Tawara, Naoyuki Kamo, Tetsuji Ogawa, Shoko Araki

arXiv 2609.39028首次发表:更新:

发表机构

1 NTT, Inc., Japan \ \ \ 2 Waseda University, Japan

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次全面分析语音增强任务的测试时优化,发现优化多个MOS预测器分数虽提升预测值,但感知质量无改善,警示评估实践需避免使用优化用预测器并保持排名预测器不公开。

AI 中文摘要

非侵入式MOS预测器被广泛用于替代主观听力测试,以评估和排名语音增强(SE)系统。如果它们能准确反映感知质量,提高其分数应能带来更高质量的语音。我们首次对SE任务的测试时优化进行了全面分析,该优化直接修改增强信号以提高多个MOS预测器分数的平均值。在URGENT 2026挑战赛的七个系统上,我们发现:1)所有优化后的预测分数均有所提高,而基于参考的指标几乎保持不变;2)未优化的预测分数没有提高;3)MUSHRA听力测试显示感知质量没有改善。这些发现揭示了一种风险,即此类优化可能扭曲评估,例如,无论感知质量如何,都会使SE系统的比较产生偏差。我们认为这些发现可以为未来的评估实践提供参考:它们表明,用于优化的预测器不应被用于评估,并且挑战赛应保持用于排名的预测器不公开。

英文摘要

Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.

Comments5 pages, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑