SLT 2026真实TSE挑战:从对话录音中提取真实世界目标说话人
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
- Nanjing University(南京大学)
- Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))
- Brno University of Technology(布拉格技术大学)
- Northwestern Polytechnical University(西北工业大学)
- NTT, Inc.(NTT公司)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
介绍SLT 2026的REAL-TSE挑战,从真实对话录音提取目标说话人,有在线和离线赛道,评估指标多样,描述了任务定义等多方面内容及经验教训。
AI中文摘要:
我们介绍了REAL-TSE挑战,这是IEEE SLT 2026关于从真实对话录音中提取目标说话人(TSE)的卫星挑战。给定多说话人混合语音和目标说话人的一个或多个注册话语,参与系统必须只恢复目标语音。与模拟朗读语音基准不同,REAL-TSE评估包含自然重叠、混响、噪声、信道失配和对话动态的普通话和英语录音。该挑战定义了两个互补赛道:用于低延迟流提取的在线赛道和用于全上下文处理的离线赛道。系统通过令牌错误率(TER)、说话人相似度(SpkSim)、DNSMOS和目标说话人活动F1进行评估。本文概述了任务定义、数据集、基线、评估协议、提交的系统、按条件的发现以及对未来真实世界TSE基准的经验教训。
英文摘要:
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.