基于互补折叠线性排序的位置训练用于多通道语音分离
Location-based Training with Complementary Folded Linear Orderings for Multichannel Speech Separation
浏览论文内容
中文总结 AI 辅助
针对多通道语音分离中圆形方位角排序的环绕不连续问题,提出基于折叠线性排序的位置训练(LBT-FLOs),利用多个互补排序的集成选择,在平面阵列和混响条件下获得稳健的适度性能提升。
中文摘要 AI 辅助
位置训练(LBT)通过施加确定性的空间排序,有效解决了多通道语音分离中的输出排列问题。对于平面麦克风阵列,LBT通常采用圆形方位角排序以覆盖整个空间范围。然而,由此产生的循环拓扑在环绕点处引入了不连续性,增加了学习复杂度并限制了对空间线索的有效利用。本研究通过引入基于折叠线性排序的位置训练(LBT-FLOs)来探究这一局限性,该方法将圆形方位角压缩为受控的线性排序。虽然单个LBT-FLO存在前后混淆,但每个LBT-FLO在特定方位角区域提供了增强的空间判别能力。利用它们的互补性,我们提出了一种集成式框架,通过方位角引导的评分在多个LBT-FLO之间进行选择。在多种平面阵列几何形状和混响条件下的实验表明,与圆形排序LBT相比,该方法取得了适度但一致的改进,并且对方位角估计误差具有鲁棒性。
英文摘要
Location-based training (LBT) effectively resolves the output permutation problem in multichannel speech separation by imposing deterministic spatial orderings. For planar microphone arrays, LBT typically adopts circular azimuth ordering to cover the full spatial range. However, the resulting cyclic topology introduces a discontinuity at the wrap-around point, increasing learning complexity and limiting the effective use of spatial cues. This work investigates this limitation by introducing location-based training with folded linear orderings (LBT-FLOs), which collapse circular azimuths into controlled linear orderings. While individual LBT-FLOs exhibit front-back ambiguity, each provides enhanced spatial discriminability over specific azimuth regions. Exploiting their complementarity, we propose an ensemble-style framework that selects among multiple LBT-FLOs using azimuth-guided scoring. Experiments across planar array geometries and reverberant conditions show modest but consistent improvements over circular-ordering LBT, with robustness to azimuth estimation errors.