DAMSEP:基于多RIR估计的距离感知单声道源分离
DAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
- Shanghai Jiao Tong University(上海交通大学)
- AISpeech Ltd(思必驰科技股份有限公司)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DAMSEP,首个联合训练用于单声道源分离和多源RIR估计的端到端框架,通过RIR的直达混响比实现距离排序,并引入HETMIXR数据集验证其优越性能。
AI中文摘要:
尽管房间冲激响应(RIRs)编码了源距离线索,但传统的单声道源分离侧重于恢复音频内容,而不估计特定源的RIRs,从而丢失了相关的空间信息。为解决这一局限,我们提出了基于多RIR估计的距离感知单声道源分离(DAMSEP),这是首个端到端框架,联合训练用于从单麦克风混合信号中进行源分离和多源RIR估计。DAMSEP集成了分离主干网络与共享的去混响和RIR估计模块,在源估计和混响重建目标下联合恢复干净源和特定源的复卷积传递函数,通过相应RIRs的直达混响比实现相对近/远排序。为进行全面评估,我们引入了HETMIXR,它涵盖异构源内容和多样化的模拟房间条件,并带有特定源RIRs和几何距离标注。在HETMIXR上的实验表明,在源分离、RIR估计和距离排序方面具有优越性能。消融研究揭示了源监督和混响重建的互补益处,而额外评估显示了对单说话者输入和使用来自未见房间的实测RIRs生成的混合信号的泛化能力。我们的代码和数据集可在该URL获取。
英文摘要:
Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture. DAMSEP integrates a separation backbone with shared dereverberation and RIR estimation modules to jointly recover clean sources and source-specific complex convolutive transfer functions under source estimation and reverberant reconstruction objectives, enabling relative near/far ordering through the direct-to-reverberant ratios of the corresponding RIRs. For comprehensive evaluation, we introduce HETMIXR, which spans heterogeneous source content and diverse simulated room conditions with source-specific RIRs and geometric distance annotations. Experiments on HETMIXR demonstrate superior performance in source separation, RIR estimation, and distance ordering. Ablation studies reveal the complementary benefits of source supervision and reverberant reconstruction, while additional evaluations show generalization to single-speaker inputs and mixtures generated using measured RIRs from an unseen room. Our code and dataset are available at https://github.com/Wenanzhi/DAMSEP.