CAL-MOS:通过适配器桥接层以实现跨语音基础模型的稳健MOS预测
CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models
- AKCIT
- Federal University of Goiás (UFG)(戈亚斯联邦大学)
- Federal University of Rio Grande do Norte (UFRN)(北里奥格兰德联邦大学)
- Federal University of Technology (UTFPR)(联邦理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出CAL-MOS方法,通过在各层池化前应用适配器进行层校准聚合,以提升跨语音基础模型的MOS预测稳健性,并缩小与完全微调的差距。
AI中文摘要:
语音质量评估(SQA)对于现代语音技术至关重要,近年来的非侵入式SQA预测器越来越依赖于语音基础模型(SFMs)。然而,由于SFMs从多个层暴露表示,目前尚不清楚哪些深度对MOS预测最具信息量,以及如何在不同的骨干网络和数据集上可靠地组合多层信息。我们在四种MOS数据集上对十种SFMs进行了基准测试,涵盖三种模式:完全微调、冻结编码器的最后一层探测以及朴素的跨层加权聚合。我们发现最佳层强烈依赖于骨干网络和数据集,并且朴素的加权融合在不同设置下可能不稳定。我们进一步评估了一种层校准的聚合变体,该变体在池化之前对每层应用适配器,这提高了多层融合的稳健性,并在保持骨干网络冻结的同时缩小了与完全微调之间的差距。
英文摘要:
Speech Quality Assessment (SQA) is essential for modern speech technologies, and recent non-intrusive SQA predictors increasingly rely on Speech Foundation Models (SFMs). However, because SFMs expose representations from many layers, it remains unclear which depths are most informative for MOS prediction and how multi-layer information should be combined reliably across backbones and datasets. We benchmark ten SFMs on four MOS datasets under three regimes: full fine-tuning, last-layer probing with a frozen encoder, and naive cross-layer weighted aggregation. We find that the best layer is strongly backbone- and dataset-dependent, and that naive weighted fusion can be unstable across settings. We further evaluate a layer-calibrated aggregation variant that applies per-layer adapters before pooling, which improves the robustness of multi-layer fusion and narrows the gap to full fine-tuning while keeping the backbone frozen.