AI 中文总结
该研究针对单声道语音提取问题,提出带跨半径一致性学习的连续半径区域语音提取方法,在实测RIR数据集上取得优于基线的性能,且对多说话人及噪声场景有效。
AI 中文摘要
声源到麦克风的距离可实现无需预注册的说话人无关语音提取。我们提出一种单声道连续半径区域语音提取方法,该方法直接对查询半径内的累积语音目标进行建模。为实现连续区域控制,查询半径被编码为连续标量并注入时频提取网络,使单个模型可在训练范围内的任意半径上运行。利用跨查询半径的目标说话人集嵌套结构,我们引入跨半径一致性学习,以稳定处理共享同一非空目标说话人集的相邻半径的预测。在实测房间冲激响应(RIR)上的实验表明,连续半径条件相比离散条件可提升选择性提取性能。所提方法取得29.25 dB的尺度不变源到失真比(SI-SDR)、4.40 dB的尺度不变源到失真比改进量(SI-SDRi)、62.01 dB的衰减量和0.46%的错误分类率(RCE),优于固定阈值和局部范围基线方法,且在说话人数量更多及存在加性噪声时仍保持有效。
英文摘要
Source-to-microphone distance enables speaker-independent speech extraction without prior enrollment. We propose a monaural continuous-radius regional speech extraction method that directly models the cumulative speech target within a queried radius. To realize continuous region control, the query radius is encoded as a continuous scalar and injected into a time-frequency extraction network, allowing a single model to operate over arbitrary radii within the trained range. Exploiting the nested structure of target-speaker sets across query radii, we introduce cross-radius consistency learning to stabilize predictions for adjacent radii sharing the same nonempty target-speaker set. Experiments on measured RIRs show that continuous-radius conditioning improves selective extraction over discrete conditioning. The proposed method achieves 29.25~dB SI-SDR, 4.40~dB SI-SDRi, 62.01~dB attenuation, and 0.46\% RCE, outperforming fixed-threshold and local-range baselines. It also remains effective with more speakers and additive noise.
CommentsSubmitted to ICASSP 2027