arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35118cs.SD

RemixIT-TSE:通过目标感知监督与重混合实现目标语音提取的渐进式合成到真实域自适应

RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing

Yu Wang, Haixin Guan, Shuang Wei, Yanhua Long

首次发表
浏览论文内容

中文总结 AI 辅助

提出RemixIT-TSE,通过两阶段渐进式合成到真实域自适应框架,结合目标感知监督与重混合技术,在真实对话场景中显著提升目标语音提取性能。

中文摘要 AI 辅助

在真实对话场景中,目标语音提取(TSE)因合成训练数据与复杂声学环境之间的域差距而遭受严重的性能下降,且此类场景中通常无法获得信号级真实标签。为应对这一挑战,我们首次尝试将RemixIT从语音增强扩展到TSE,并提出一种用于真实世界TSE的渐进式合成到真实域自适应框架,包含两个微调阶段。第一阶段利用区域级说话人相似性和静音约束,在目标感知自适应框架内,联合使用合成数据和弱监督真实数据优化模型,在注入真实世界特征的同时保留合成数据学习到的能力。第二阶段仅使用真实数据,通过我们的RemixIT-TSE进一步适配模型,其中经质量过滤的教师伪目标(确保可靠的教师训练)通过SI-SNR损失提供信号级监督。在SLT 2026 REAL-TSE挑战赛的真实对话评估集(EVAL-2)上,所提方法相比源域基线实现了6.53%的相对词错误率(TER)降低,同时在说话人相似性、DNSMOS-P808和目标活动F1上分别获得21.84%、9.89%和4.10%的相对提升,证明了其在未见过的真实条件下的有效性。源代码见https://github.com/YuWang-Speech/RemixIT-TSE。

英文摘要

Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.

发表机构

  • Shanghai Normal University(上海师范大学)
  • Unisound AI Technology Co., Ltd.(云知声智能科技股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑