arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于掩码的语音增强在空间音频中的应用:Ambisonics、波束成形与麦克风通道的比较

Mask-Based Speech Enhancement for Spatial Audio: A Comparison of Ambisonics, Beamforming, and Microphone Channels

Sheli Hendel, Boaz Rafaely, Dorothea Kolossa

arXiv 2609.18532首次发表:更新:

AI 中文总结

本研究系统比较了掩码语音增强在麦克风、波束成形和Ambisonics三种空间音频表示中的性能,发现波束成形域增强效果最佳,而Ambisonics域更利于保留空间属性。

AI 中文摘要

基于掩码的语音增强被广泛用于抑制噪声和干扰,但其在具有多通道输出的空间音频算法中的性能尚未得到广泛研究。在此类设置中,语音增强必须在提升语音质量的同时保留对定位、空间感知和空间释放掩蔽至关重要的空间线索。在本工作中,我们系统地比较了应用于三种信号表示的时间-频率掩码:麦克风信号、波束成形器输出和Ambisonics信号。性能评估涵盖语音质量、可懂度、双耳线索保留和混响保留。结果揭示了增强与空间保真度之间的明确权衡:波束成形域掩码实现了最高的语音增强得分,而Ambisonics域掩码更好地保留了残余干扰的空间属性。所有方法均保留了目标的定位线索。

英文摘要

Mask-based speech enhancement is widely used for suppressing noise and interference, but its performance in spatial audio algorithms with multichannel output has not been studied extensively. In such settings, speech enhancement must improve speech quality while preserving spatial cues that are essential for localization, spatial awareness, and spatial release from masking. In this work, we systematically compare time frequency masking applied to three signal representations: microphone signals, beamformer outputs, and Ambisonics signals. Performance is evaluated in terms of speech quality, intelligibility, binaural cue preservation, and reverberation preservation. Results reveal a clear trade-off between enhancement and spatial fidelity: beamformer-domain masking achieves the highest speech enhancement scores, while Ambisonics-domain masking better preserves the spatial attributes of the residual interference. All methods preserve the target's localization cues.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑