arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11092eess.AS

下游任务感知的统一源分离

Downstream-Task-Aware Unified Source Separation

Yoshiki Mitsui, Ryo Aihara, Tatsuhiko Saito, Yoshiki Masuyama, Christoph Boeddeker, Julius Richter, Gordon Wichern, Jonathan Le Roux

AI总结:

本文提出一种下游任务感知的提示扩展框架,将任务信息融入TUSS提示并切换损失函数,使单一模型既能提升ASR性能又能保持语音增强质量。

AI中文摘要:

任务感知的统一源分离(TUSS)通过输入提示条件化,使单一模型能够处理多种分离任务。然而,传统TUSS不考虑下游任务需求,例如增强后的语音是用于人工聆听还是自动语音识别(ASR)。本文提出了一种TUSS的提示扩展框架,将下游任务信息纳入输入提示,并在训练过程中根据给定提示切换损失函数,从而在推理时输出具有不同信号特征的语音。具体而言,我们引入了一个专用于ASR的提示,并配以正则化损失函数,以减少语音伪影、提升ASR鲁棒性;而标准提示则配以传统的信噪比(SNR)损失函数。在LibriSpeech和JNAS语料库上的实验表明,所提出的联合训练方案使单一模型能够通过选择ASR专用提示,在广泛的SNR条件下提升ASR性能(相对于含噪输入),同时在使用标准提示时保持通用语音增强质量。

英文摘要:

Task-aware unified source separation (TUSS) enables a single model to handle diverse separation tasks by conditioning on input prompts. However, conventional TUSS does not account for downstream task requirements, such as whether the enhanced speech will be used for human listening or automatic speech recognition (ASR). In this paper, we propose a prompt extension framework for TUSS that incorporates downstream task information into the input prompts and switches the loss function according to the given prompt during training, enabling outputs with different signal characteristics at inference time. Specifically, we introduce an ASR-dedicated prompt paired with a regularized loss function that reduces speech artifacts to improve ASR robustness, while the standard prompt is paired with the conventional SNR loss function. Experiments on the LibriSpeech and JNAS corpora demonstrate that the proposed joint-training scheme enables a single model to improve ASR performance over noisy input across a wide range of SNR conditions by selecting the ASR-dedicated prompt, while maintaining general speech enhancement quality when the standard prompt is used.

补充信息

↑