arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.02571eess.AScs.AIcs.LGcs.SDeess.SP

基于物理意识的深度复合核的转向矢量高斯过程回归用于增强听觉

Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening

  • National Institute of Informatics(国家信息研究所)

机构由 AI 辅助整理,请以论文原文为准。

Diego Di Carlo, Shoichi Koyama, Nugraha Aditya Arie, Fontaine Mathieu, Bando Yoshiaki, Yoshii Kazuyoshi

更新

AI总结:

本文提出基于神经场的高斯过程框架,结合物理意识复合核,用于增强听觉中转向矢量的连续表示,以实现更有效的空间滤波和双耳渲染。

AI中文摘要:

本文研究了在频率和麦克风/声源位置上连续表示转向矢量以实现增强听觉(如空间滤波和双耳渲染)的连续表示,从而实现用户参数化的再现声音场控制。转向矢量通常用于表示麦克风阵列的空间响应作为仰角函数。假设理想环境的基本代数表示无法处理声场的散射效应。因此,可以收集在专用设施中测量的离散转向矢量集合并进行超分辨率(即上采样)。最近,物理意识深度学习方法已被有效用于此目的。然而,这种确定性超分辨率由于测量空间中的非均匀不确定性而面临过拟合问题。为了解决这个问题,我们将基于神经场(NF)的表达性表示整合到基于高斯过程(GP)的原理性概率框架中。具体来说,我们提出了一种物理意识复合核,用于建模入射波的方向和随后的散射效应。我们的综合比较实验表明,在数据不足条件下所提方法的有效性。在下游任务如使用SPEAR挑战模拟数据的语音增强和双耳渲染中,通过少于十倍的测量数据获得了最优性能。

英文摘要:

This paper investigates continuous representations of steering vectors over frequency and microphone/source positions for augmented listening (e.g., spatial filtering and binaural rendering), enabling user-parameterized control of the reproduced sound field. Steering vectors have typically been used for representing the spatial response of a microphone array as a function of the look-up direction. The basic algebraic representation of these quantities assuming an idealized environment cannot deal with the scattering effect of the sound field. One may thus collect a discrete set of real steering vectors measured in dedicated facilities and super-resolve (i.e., upsample) them. Recently, physics-aware deep learning methods have been effectively used for this purpose. Such deterministic super-resolution, however, suffers from the overfitting problem due to the non-uniform uncertainty over the measurement space. To solve this problem, we integrate an expressive representation based on the neural field (NF) into the principled probabilistic framework based on the Gaussian process (GP). Specifically, we propose a physics-aware composite kernel that models the directional incoming waves and the subsequent scattering effect. Our comprehensive comparative experiment showed the effectiveness of the proposed method under data insufficiency conditions. In downstream tasks such as speech enhancement and binaural rendering using the simulated data of the SPEAR challenge, the oracle performances were attained with less than ten times fewer measurements.

↑