arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProxyMOS:通过多教师蒸馏与自适应路由实现无标签语音质量评估

ProxyMOS: Label-Free Speech Quality Assessment by Multi-Teacher Distillation with Adaptive Routing

Maxim Trokunov, Kirill Borodin, Nikita Vasiliev, Grach Mkrtchian

arXiv 2610.00419首次发表:更新:

发表机构

Lab260; BitmanagerAI; MTUCI(Lab260; BitmanagerAI; 莫斯科通信与信息技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ProxyMOS通过多教师蒸馏与自适应路由,将多个公开MOS预测器集成并训练学生模型,无需人工标签,在URGENT和mos260基准上超越最佳教师,实现高效无标签语音质量评估。

AI 中文摘要

人类平均意见得分(MOS)的收集成本高昂,且非侵入式MOS预测器在其训练领域之外性能急剧下降。ProxyMOS将一组公开的MOS预测器转化为一个更强的单一模型,无需新的人类标签。八个预测器与人类评分进行基准测试;五个信息量最大的预测器进入子集搜索,采用均匀、相关性加权、误差加权、均方误差优化和自适应逐话语路由;最佳路由的四模型集成对807k条未标记话语进行标注,用于训练一个wav2vec 2.0学生模型。在URGENT上,学生模型达到Spearman ρ=0.802,而最佳教师为0.773。在mos260(一个新的俄语TTS基准,包含来自38个合成条件的4,600条话语)上,它达到ρ=0.636(逐话语)和0.95(逐条件),在单次前向传播中匹配其自身的路由集成。自适应路由是唯一在添加弱预测器时不降级的规则。模型、ONNX导出和mos260已发布。

英文摘要

Human mean opinion scores (MOS) are costly to collect, and non-intrusive MOS predictors degrade sharply outside their training domain. ProxyMOS turns a pool of public MOS predictors into a single stronger model without new human labels. Eight predictors are benchmarked against human ratings; the five most informative enter a subset search under uniform, correlation-weighted, error-weighted, MSE-optimised and adaptive per-utterance routing; and the best routed four-model ensemble labels 807k unlabeled utterances that train a wav2vec 2.0 student. On URGENT the student reaches Spearman $ρ=0.802$ against $0.773$ for the best teacher. On mos260, a new Russian TTS benchmark of 4,600 utterances from 38 synthesis conditions, it reaches $ρ=0.636$ against $0.613$ per utterance and $0.95$ per condition, matching its own routed ensemble in one forward pass. Adaptive routing is the only rule that does not degrade when weak predictors are added. Model, ONNX exports and mos260 are released. It's about 950 characters; arXiv's limit is 1,920. I kept $ρ$ because arXiv renders it on the abstract page. If you'd rather avoid math, replace $ρ=0.802$ with rho = 0.802 and do the same for the other $...$ values.

CommentsSubmitted to IEEE ICASSP 2027. 5 pages + 7 pages supplementary material. Model: https://huggingface.co/lab260/ProxyMos, benchmark: https://huggingface.co/datasets/lab260/mos260, code: https://github.com/lab260ru/ProxyMOS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑