arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过留一数据集验证评估用于跨语料库音频MOS预测的SSL和ViViT架构

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

Mustafa Ozan Duman, Ahmet Emir Dirik

arXiv 2607.10146首次发表:更新:

AI 中文总结

研究针对跨语料库音频MOS预测,对SSL-FRZ、SSL-FT和ViViT三种架构进行基准测试,采用留一数据集验证评估稳健性,用ARECHO框架对比18个指标,发现纯英文语料库精度更高,SSL-FRZ在不可见分布上稳健性好,为语音质量评估提供稳定可扩展方案。

AI 中文摘要

自动平均意见得分(MOS)预测对于评估大规模合成语音和音频增强系统至关重要,但模型常受域转移困扰。本研究对三种架构框架进行了全面基准测试:冻结自监督学习(SSL-FRZ)、微调SSL(SSL-FT)和视频视觉Transformer(ViViT)。分两阶段评估,第一阶段使用19个不同数据集的130,000个样本的合并语料库,第二阶段聚焦纯英文的17数据集语料库。采用留一数据集(LODO)协议评估稳健性,最后用ARECHO框架根据18个最新指标对最佳模型进行基准测试。结果表明,纯英文语料库在所有架构中预测精度更高。SSL-FT在可见验证数据上性能最高,SSL-FRZ在不可见分布上稳健性更好,ViViT在纯英文测试中结果稳定。LODO结果证实,冻结SSL嵌入与深度Transformer编码器结合为通用语音质量评估提供了最稳定且可扩展的解决方案。为支持进一步研究,最佳纯英文SSL-Transformer模型和权重通过Hugging Face公开。

英文摘要

Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. This study presents a comprehensive benchmarking of three architectural frameworks: Frozen Self-Supervised Learning (SSL-FRZ), Fine-Tuned SSL (SSL-FT), and a Video Vision Transformer (ViViT). Evaluation is conducted in two phases: Part I utilizes a consolidated corpus of 130,000 samples across 19 diverse datasets, while Part II focuses on a purified 17-dataset English-only corpus. To assess robustness, a systematic Leave-One-Dataset-Out (LODO) protocol is employed to quantify the generalization gap between seen and unseen distributions. Finally, the top-performing model is benchmarked against 18 state-of-the-art (SOTA) metrics using the ARECHO framework. Results demonstrate that an English-only purified corpus consistently yields higher predictive precision across all architectures. While SSL-FT achieves the highest performance on seen validation data, the SSL-FRZ model provides superior robustness on unseen distributions, achieving a competitive Mean Squared Error (MSE) of 0.36 on the URGENT 2024 benchmark-closely matching domain-optimized SOTA metrics (MSE 0.30). Although the ViViT architecture remains below SSL-based models in total capacity, it delivers stable results in English-only trials. LODO results confirm that while models perform significantly better on seen samples, frozen SSL embeddings combined with deep Transformer encoders offer the most stable and scalable solution for universal speech quality assessment. To support further research, the top-performing English-only SSL-Transformer model and weights are made publicly available via Hugging Face.

CommentsHuggingface link: https://huggingface.co/mustafa-ozan-duman/wavlm-transformer-mos-english Github link: https://github.com/mustafa-ozan/audio_mos_prediction_SSL_ViViT_codes

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑