arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DAMOS:通过显式失真定位学习感知失真的语音质量评估

DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization

Naiyuan Li, Li Dong, Diqun Yan

arXiv 2608.21176首次发表:更新:

发表机构

College of Information Science and Engineering, Ningbo University; College of Artificial Intelligence, Ningbo University of Finance and Economics(宁波大学信息科学与工程学院; 宁波财经学院人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有语音质量评估方法无法明确指示失真位置的局限,构建带帧级标注的数据集,提出DAMOS框架,通过显式失真定位辅助MOS预测,在多基准上优于现有方法且泛化性强。

AI 中文摘要

自动语音质量评估旨在预测与人类主观感知一致的平均意见得分(MOS),对评估语音生成、增强和通信系统至关重要。对于语音信号,尤其是合成语音,失真通常是局部发生的,整体感知质量通常由少数感知显著的失真区域主导。然而,大多数现有方法主要针对语句级MOS进行优化,仅提供粗粒度监督,无法明确指示感知重要失真发生的位置。为解决这一局限,我们引入显式失真定位作为语音质量评估的辅助知识,构建了首个具有帧级失真标注的部分失真语音数据集,并训练定位模型生成失真线索。基于这些线索,我们提出了DAMOS,这是一种感知失真的语音质量评估框架,将定位信息整合到MOS预测流程中。在多个公开基准上的实验表明,DAMOS始终优于现有方法,且表现出强大的跨数据集泛化能力,验证了显式失真定位对语音质量评估的有效性。

英文摘要

Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evaluating speech generation, enhancement, and communication systems. For speech signals, especially synthetic speech, distortions often occur locally, and overall perceptual quality is usually dominated by a small number of perceptually salient distortion regions. However, most existing methods are primarily optimized with utterance-level MOS, which provides only coarse-grained supervision and offer no explicit indication of where perceptually important distortions occur. To address this limitation, we introduce explicit distortion localization as auxiliary knowledge for speech quality assessment. We construct the first partially distorted speech dataset with frame-level distortion annotations and train a localization model to generate distortion cues. Building on these cues, we propose DAMOS, a distortion-aware speech quality assessment framework that integrates localization information into the MOS prediction pipeline. Experiments on multiple public benchmarks demonstrate that DAMOS consistently outperforms existing methods and exhibits strong cross-dataset generalization, validating the effectiveness of explicit distortion localization for speech quality assessment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑