arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27436eess.AS

WeSep:面向目标说话人提取的模块化与线索可组合框架

WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

Ke Zhang, Xiaoyang Yu, Haoyu Li, Shuai Wang, Shuhan Zhang, Haizhou Li

首次发表
浏览论文内容

中文总结 AI 辅助

WeSep是将目标说话人提取重新表述为异质线索条件学习问题的统一框架,通过模块化设计提升灵活性,多模态实验验证其稳定性,工具包将公开。

中文摘要 AI 辅助

目标说话人提取(TSE)旨在借助辅助线索从重叠语音混合中分离出目标说话人。现有系统通常针对特定线索类型设计,当不同场景线索可用性变化时灵活性受限。本文提出WeSep,一种将TSE重新表述为异质线索条件学习问题的统一框架。WeSep中,线索模块与分离骨干通过标准化接口解耦,支持可配置线索注入及多模态灵活集成。该设计可在共享优化框架内系统研究线索结构、模态内与跨模态交互、动态线索可用性,助力适配真实场景。针对注册、空间、视觉及文本线索的实验揭示了模态依赖特性,并证明异质线索可用性下优化稳定,该工具包将公开提供。

英文摘要

The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSep, a unified framework that reformulates TSE as a heterogeneous cue-conditioned learning problem. In WeSep, cue modules and separator backbones are decoupled through standardized interfaces, enabling configurable cue injection and flexible integration of diverse modalities. The design enables systematic study of cue structure, intra- and cross-modal interaction, and dynamic cue availability within a shared optimization framework, facilitating adaptation to real-world conditions. Experiments across enrollment, spatial, visual, and textual cues reveal modality-dependent characteristics and demonstrate stable optimization under heterogeneous cue availability. The toolkit will be publicly available.

补充信息

↑