arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15198eess.AScs.SD

SLT 2026真实TSE挑战:从对话录音中提取真实世界目标说话人

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

  • Nanjing University(南京大学)
  • Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))
  • Brno University of Technology(布拉格技术大学)
  • Northwestern Polytechnical University(西北工业大学)
  • NTT, Inc.(NTT公司)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han, Zikai Liu, Xiaoyang Yu, Haoyu Li, Marc Delcroix, Kai Yu, Lei Xie, Ming Li, Haizhou Li

AI总结:

介绍SLT 2026的REAL-TSE挑战,从真实对话录音提取目标说话人,有在线和离线赛道,评估指标多样,描述了任务定义等多方面内容及经验教训。

AI中文摘要:

我们介绍了REAL-TSE挑战,这是IEEE SLT 2026关于从真实对话录音中提取目标说话人(TSE)的卫星挑战。给定多说话人混合语音和目标说话人的一个或多个注册话语,参与系统必须只恢复目标语音。与模拟朗读语音基准不同,REAL-TSE评估包含自然重叠、混响、噪声、信道失配和对话动态的普通话和英语录音。该挑战定义了两个互补赛道:用于低延迟流提取的在线赛道和用于全上下文处理的离线赛道。系统通过令牌错误率(TER)、说话人相似度(SpkSim)、DNSMOS和目标说话人活动F1进行评估。本文概述了任务定义、数据集、基线、评估协议、提交的系统、按条件的发现以及对未来真实世界TSE基准的经验教训。

英文摘要:

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.

补充信息

↑