arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38000eess.AS

QK-GCC:可学习的查询-键频谱匹配用于鲁棒时延估计

QK-GCC: Learnable Query-Key Spectral Matching for Robust Time Delay Estimation

Jinkai Zhang, Weiye Chen, Yue Huang, Xiaotong Tu, Xinghao Ding

首次发表
浏览论文内容

中文总结 AI 辅助

提出QK-GCC,一种可学习的类GCC框架,用查询-键匹配替代手工频谱匹配,提升时延估计在噪声和混响下的鲁棒性,优于GCC-PHAT及学习变体。

中文摘要 AI 辅助

时延估计(TDE)是麦克风阵列声源定位的基本组成部分。广义互相关(GCC)因其高效性和可解释性而被广泛使用,但其手工设计的频谱匹配和预定义的频率加权易受噪声和混响的影响。现有的神经GCC变体主要通过增强输入信号或建模GCC响应来提高鲁棒性,而跨通道频谱匹配步骤本身仍然是手工设计的。我们提出QK-GCC,一种可学习的类GCC框架,用两个麦克风信号之间的查询-键(Query-Key)匹配取代GCC中手工设计的加权频谱匹配。两个麦克风信号被编码为幅度-相位频率令牌,并分别映射到查询(Query)和键(Key)表示,从而实现频率可靠性学习和局部频谱证据聚合以进行时延估计。在模拟混响房间中跨不同信噪比(SNR)和混响条件下的实验表明,QK-GCC在时延估计精度上优于GCC-PHAT和基于学习的GCC变体,同时保持轻量级并泛化到未见过的声源类型。代码可在该https URL获取。

英文摘要

Time delay estimation (TDE) is a fundamental component of microphone-array sound source localization. Generalized cross-correlation (GCC) is widely used because it is efficient and interpretable, but its handcrafted spectral matching and predefined frequency weighting are vulnerable to noise and reverberation. Existing neural GCC variants mainly improve robustness by enhancing input signals or modeling GCC responses, while the cross-channel spectral matching step itself remains handcrafted. We propose QK-GCC, a learnable GCC-like framework that replaces handcrafted weighted spectral matching in GCC with Query-Key matching between two microphone signals. The two microphone signals are encoded as magnitude-phase frequency tokens and mapped to Query and Key representations, respectively, enabling frequency reliability learning and local spectral evidence aggregation for delay estimation. Experiments in simulated reverberant rooms across diverse SNR and reverberation conditions show that QK-GCC improves TDE accuracy over GCC-PHAT and learning-based GCC variants, while remaining lightweight and generalizing to unseen source types. The code is available at https://github.com/zhangjinkai33-ui/QK-GCC.

补充信息

↑