AI 中文总结
针对人机交互中重叠声音时序识别的难题,提出多分支CNN结合注意力融合的深度学习框架,在三类实验条件下取得较高准确率,验证了实时可行性。
AI 中文摘要
识别重叠声音的时间顺序是人机交互(HRI)中一个研究不足的挑战,与急救人员检测系统等应用直接相关。本文提出一种用于实时声音时序识别的深度学习框架,该框架使用可录制的蜂鸣器发出不同的非语言声音(猫叫、狗吠、直升机噪声)。多分支卷积神经网络(CNN)处理梅尔频谱图、梅尔频率倒谱系数(MFCC)和短时傅里叶变换(STFT)特征,采用基于注意力的融合机制来强调关键时间线索。实验在等振幅、变振幅和未见过的声音条件下开展,所提系统在平衡重叠情况下准确率达99%,振幅变化下为91%,经归一化处理的未见过测试数据上为74%。这些结果表明,深度学习可在重叠条件下可靠识别声音时序,支持实际人机交互场景。尽管实验在精心控制的合成重叠数据上进行,我们还报告了延迟基准测试以证明实时可行性,并就真实房间环境中的泛化、生态效度和部署挑战展开了拓展讨论。
英文摘要
Recognizing the temporal order of overlapping sounds is an underexplored challenge in human-robot interaction (HRI), with direct relevance to applications such as first responder detection systems. This paper presents a deep learning framework for real-time sound order recognition using recordable buzzers that emit distinct non-verbal sounds (cat meows, dog barks, helicopter noises). A multi-branch convolutional neural network (CNN) processes Mel spectrograms, Mel-frequency cepstral coefficients (MFCCs), and short-time Fourier transform (STFT) features, with an attention-based fusion mechanism to emphasize critical temporal cues. Experiments were conducted under same-amplitude, varied-amplitude, and unseen sound conditions. The proposed system achieved 99% accuracy in balanced overlaps, 91% under amplitude variation, and 74% on unseen test data with normalization. These results demonstrate that deep learning can reliably recognize sound order in overlapping conditions, supporting practical HRI scenarios. While experiments were conducted on carefully controlled synthetic overlaps, we additionally report latency benchmarks demonstrating real-time feasibility and provide an extended discussion on generalization, ecological validity, and deployment challenges in real-room environments.