arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

口语语言识别中的声学域偏移:从系统域泛化评估到实际应用

Acoustic domain shift in spoken language identification from systematic domain generalization evaluation to real-world application

Francois Derrida, Raphaël Duroselle, Thomas Courtat, Jean-François Bonastre

arXiv 2609.31759首次发表:更新:

发表机构

THALES(泰雷兹集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出小规模语音数据集和评估协议,系统研究口语语言识别中的声学域偏移,发现域内性能不能预测跨域鲁棒性,且显式域不变算法不优于ERM,并在MMS-LID-126上验证。

AI 中文摘要

域泛化(DG)旨在开发对训练期间未见条件保持鲁棒的模型。虽然域泛化已通过受控基准和多样分布偏移在计算机视觉中进行了系统研究,但其在口语语言识别中的评估仍缺乏结构化。现有语音数据集为对真实世界声学条件的鲁棒性提供了有价值的基准,但主要围绕特定场景和大规模设计,而非作为系统构建和评估域偏移的通用工具。在本工作中,我们引入一个小规模语音数据集和评估协议,用于声学域偏移的受控研究。它能够在多种声学域偏移下评估口语语言识别模型。我们将语音模态引入DomainBed域泛化框架,并评估三种域泛化算法。我们表明,域内性能不是跨域鲁棒性的可靠预测指标,并验证了诸如MMD或DANN等显式域不变算法并不优于ERM。我们进一步在MMS-LID-126(一种最先进的口语语言识别系统)上验证了这些发现的普遍性。我们发布了代码。

英文摘要

Domain Generalization (DG) aims to develop models that remain robust to conditions unseen during training. While DG has been systematically studied in computer vision through controlled benchmarks and diverse distribution shifts, its evaluation in spoken language recognition remains less structured. Existing speech datasets provide valuable benchmarks for robustness to real-world acoustic conditions, but are primarily designed around specific scenarios and large scale rather than as general-purpose tools for systematically constructing and evaluating domain shifts. In this work, we introduce a smallscale speech dataset and evaluation protocol for controlled studies of acoustic domain shifts. It enables the evaluation of spoken language recognition models under a variety of acousitc domain shifts. We introduce the speech modality into the DomainBed Domain Generalization framework and evaluate three domain generalization algorithms. We show that in-domain performance is not a reliable predictor of cross-domain robustness and verify that explicit domain invariant algorithms such as MMD or DANN algorithms do not outperform ERM. We further validate the generality of these findings on MMS-LID-126, a state-of-the-art spoken language identification system. We release the code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑