采用物理启发式扩散模型的可变拓扑稀疏麦克风阵列的几何自适应Ambisonic编码
Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion
浏览论文内容
中文总结 AI 辅助
本文提出DiffM2A几何自适应条件扩散框架,结合GASHP前端与双分支阐释扩散模型,在稀疏可变拓扑麦克风阵列的Ambisonic编码任务中,优于基线方法且在未知布局与边界模型下增益稳定。
中文摘要 AI 辅助
Ambisonics提供紧凑的基于场景的空间音频表示,但高阶Ambisonic编码对可穿戴设备和嵌入式硬件构成困难。这些设备的麦克风阵列通常是稀疏、不规则的,且受设备特定边界条件限制。这些因素使得球谐(SH)域编码病态:逆滤波会放大噪声,而确定性神经编码器可能过拟合于阵列特定响应或平滑模糊的高阶分量。本文提出DiffM2A,一种用于从具有可变拓扑的稀疏麦克风阵列(MAs)进行稳健Ambisonic编码的几何自适应条件扩散框架。其几何自适应球谐投影(GASHP)前端构建感知边界的SH导向函数,并应用能量归一化模态投影,将阵列相关观测映射到公共模态表示而无需显式伪逆计算。双分支阐释扩散模型随后估计复Ambisonic系数,以原始麦克风频谱和GASHP特征为条件。声强和旋转等变损失进一步增强跨SH子空间的通道间相位一致性和结构化行为。对一阶和二阶Ambisonic编码任务的评估,使用模拟房间声学和真实LOCATA录音,表明DiffM2A在信号保真度、频谱精度、空间相干性和双耳线索保留方面优于传统和神经基线方法。额外实验显示,这些增益在未见的五麦克风布局以及不匹配的开阵列和刚性球边界模型下基本得以保留。
英文摘要
Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.
发表机构
- Center of Intelligent Acoustics and Immersive Communications, School of Artificial Intelligence, Northwestern Polytechnical University(西北工业大学人工智能学院智能声学与沉浸式通信中心)
- Digital Signal Processing Lab, School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院数字信号处理实验室)
机构由 AI 辅助整理,请以论文原文为准。