发表机构
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Beijing University of Posts and Telecommunications; Dalian Maritime University(中国科学院自动化研究所; 中国科学院大学; 北京邮电大学; 大连海事大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RadarVox提出雷达-音频多模态基准,利用FMCW雷达捕获喉部运动线索,通过门控融合和跨模态匹配实现身份感知的鸡尾酒会语音分离,显著提升分离质量和说话人分配准确率。
AI 中文摘要
在具身语音交互中,鸡尾酒会语音感知需要同时实现语音分离和说话人归属,以区分机器生成语音与人类语音源。然而,传统的纯音频盲源分离存在排列模糊性问题,使得分离流与物理说话人之间的对应关系不明确。本文提出RadarVox,一个用于身份感知鸡尾酒会语音分离的雷达-音频多模态基准。RadarVox提供来自扬声器发射和人类语音的声学混合信号及源级雷达位移信号,实现源感知的说话人分配。FMCW雷达捕获喉部机械运动,提供单通道麦克风无法获得的说话人特定空间运动线索。我们通过门控融合将雷达衍生的先验注入DPRNN分离器,并学习说话人感知的跨模态匹配器,将无序语音流与雷达观测到的说话人关联。在多说话人混合实验表明,雷达线索提供了互补优势,在双说话人和三说话人场景中分别实现了9.75 dB和7.13 dB的尺度不变信号失真比(SI-SDR)。更重要的是,所提方法在有序SI-SDR上比纯音频方法提高了超过13 dB,同时实现了超过80%的说话人分配准确率。
英文摘要
In embodied voice interaction, cocktail-party speech perception requires both speech separation and speaker attribution across machine-generated and human speech sources. However, conventional audio-only blind source separation remains permutation ambiguous, making the correspondence between separated streams and physical speakers unclear. This paper presents RadarVox, a radar-audio multimodal benchmark for identity-aware cocktail-party speech separation. RadarVox provides acoustic mixtures and source-level radar displacement signals from loudspeaker-emitted and human speech, enabling source-aware speaker assignment. FMCW radar captures laryngeal mechanical motion, providing speaker-specific spatial-motion cues unavailable to a single-channel microphone. We inject radar-derived priors into a DPRNN separator via gated fusion and learn a speaker-aware cross-modal matcher to associate unordered speech streams with radar-observed speakers. Experiments on multi-speaker mixtures show that radar cues provide complementary benefits, achieving scale-invariant signal-to-distortion ratio (SI-SDR) values of 9.75 dB and 7.13 dB in two- and three-speaker scenarios, respectively. More importantly, the proposed method improves ordered SI-SDR by more than 13 dB over audio-only methods while achieving over 80\% speaker assignment accuracy.