发表机构
Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua University; IDG/McGovern Institute for Brain Research, Tsinghua University; Chinese Institute for Brain Research (CIBR)(计算机科学与技术系、人工智能研究所、BNRist、清华大学; IDG/麦戈文脑科学研究所、清华大学; 中国脑科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍 SonicAGI 参加 REAL-TSE 挑战赛的情况,采用以数据为中心的方法,在线引入 SwiftNet-Lookahead 控制延迟,离线用 USEF-TFGridNet 权衡质量与保真度,在官方评估中取得较好成绩,证明相关训练和建模对对话式 TSE 有效。
AI 中文摘要
现实世界中的目标说话人提取(TSE)仍然具有挑战性,因为目标语音、干扰和注册是在声学条件不匹配的情况下录制的,存在混响、噪声和不规则的对话重叠。本文描述了 SonicAGI 提交给 REAL-TSE 挑战赛(IEEE SLT 2026)的情况。我们采用以数据为中心的方法,将干净语音的完全模拟混合与真实会议重叠相结合,并使用冻结的离线增强器为辅助监督提供真实目标的去噪镜像。对于在线赛道,我们引入了 SwiftNet-Lookahead,它在严格因果迭代分离器之前插入一个有界前瞻模块,并将总系统延迟保持在 96 毫秒。对于离线赛道,我们使用具有幅度域融合阶段的帧级注册交叉注意力 USEF-TFGridNet,该阶段在感知质量和说话人保真度之间进行权衡。在官方评估中,SwiftNet-Lookahead 在赛道 1 中排名第二,USEF-TFGridNet 在赛道 2 中排名第五,均超过了挑战赛基线。这些结果表明,面向真实数据的训练和特定赛道的建模对于对话式 TSE 是有效的。
英文摘要
Real-world target speaker extraction (TSE) remains challenging because target speech, interference, and enrollment are recorded under mismatched acoustic conditions with reverberation, noise, and irregular conversational overlap. This paper describes the SonicAGI submission to the REAL-TSE Challenge (IEEE SLT 2026). We take a data-centric approach that combines fully simulated mixtures from clean speech with real meeting overlaps, and use a frozen offline enhancer to provide a denoised mirror of real targets for auxiliary supervision. For the online track, we introduce SwiftNet-Lookahead, which inserts a single bounded-lookahead module before a strictly causal iterative separator and keeps the total system latency at 96 ms. For the offline track, we use a frame-level enrollment cross-attention USEF-TFGridNet with a magnitude-domain fusion stage that trades off perceptual quality and speaker fidelity. In the official evaluation, SwiftNet-Lookahead ranks second in Track~1 and USEF-TFGridNet ranks fifth in Track~2, both exceeding the challenge baselines. These results suggest that real-data-oriented training and track-specific modeling are effective for conversational TSE.