arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31787eess.AScs.AIcs.CLcs.LG

最优传输与语音的相遇:一篇教程式综述

Optimal transport meets speech: a tutorial review

Xugang Lu, Yu Tsao

首次发表
浏览论文内容

中文总结 AI 辅助

本文是一篇教程式综述,系统介绍最优传输(OT)理论及其在语音处理中的应用,涵盖基础原理、深度学习算法及跨域跨模态任务(如语音增强、识别与防欺骗),旨在推动OT在语音领域的更广泛采用。

中文摘要 AI 辅助

最优传输(Optimal Transport, OT)为比较和变换概率分布提供了一个有原则的框架,同时保持几何结构。近年来,OT在机器学习领域获得了显著关注,因为它能够度量分布之间的差异,即使这些分布的支撑集不重叠,这使得它在生成建模、领域自适应和迁移学习等任务中非常有效。尽管OT在计算机视觉和自然语言处理等领域取得了成功,但在语音研究中仍相对未被充分探索。语音信号带来了独特的挑战,包括时间动态、说话人变异性、噪声、混响以及涉及音频、文本和视觉信息的异构多模态表示。这些因素常常导致分布不匹配,而OT为对齐和解释提供了一个自然的框架。本工作旨在促进OT在语音处理中的更广泛采用,具体方式包括:(1)通过直观的物理解释回顾OT基础,并强调其与现代生成模型的联系;(2)介绍适用于深度学习框架的计算算法;(3)展示OT在跨领域和跨模态语音任务中的应用,包括语音增强、自动语音识别、语言和说话人识别以及音频欺骗检测。我们强调OT在解决真实世界语音应用中的分布变化方面具有巨大潜力。

英文摘要

Optimal Transport (OT) provides a principled framework for comparing and transforming probability distributions while preserving geometric structure. Recently, OT has gained significant attention in machine learning due to its ability to measure discrepancies between distributions, even when their supports do not overlap, making it effective for tasks such as generative modeling, domain adaptation, and transfer learning. Despite its success in fields such as computer vision and natural language processing, OT remains relatively underexplored in speech research. Speech signals present unique challenges, including temporal dynamics, speaker variability, noise, reverberation, and heterogeneous multimodal representations involving audio, text, and visual information. These factors often lead to distribution mismatches, where OT offers a natural framework for alignment and interpretation. This work aims to promote broader adoption of OT in speech processing by: (1) reviewing OT foundations through intuitive physical interpretations and highlighting connections to modern generative models; (2) presenting computational algorithms suitable for deep learning frameworks; and (3) demonstrating OT applications in cross-domain and cross-modal speech tasks, including speech enhancement, automatic speech recognition, language and speaker recognition, and audio spoof detection. We highlight OT's strong potential for addressing distributional variations in real-world speech applications.

发表机构

  • National Institute of Information and Communications Technology(国立信息与通信技术研究所)
  • Research Center for Information Technology Innovation, Academia Sinica(中央研究院信息科技创新研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑