发表机构
Reality Labs Research, Meta; Aalto University(元宇宙实验室研究部,Meta; 阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对不规则/稀疏等空间捕获能力有限的麦克风阵列无法用线性方法重建HOA RIR高阶空间细节的问题,提出基于扩散的生成框架,实现设备无关编码,实验显示其性能优于基线,可支撑可扩展声学模拟。
AI 中文摘要
我们解决了将房间冲激响应(RIR)编码为高阶Ambisonics(HOA)表示的问题,该编码基于任意且可能不足或不完整的麦克风阵列测量。对于空间捕获能力有限的麦克风阵列,如不规则或稀疏阵列,此任务本质上是不适定的,因为经典线性方法无法重建高阶空间细节。我们引入了一种基于扩散的生成框架,该框架对HOA RIR的统计特性进行建模,从而能够从任意麦克风阵列(可能在数据测量期间未见过)进行设备无关编码。我们的方法包含后验采样过程,该过程在估计信号与测量之间保持一致性,同时合理重建仅从有限测量中不可观察的空间信息。对模拟数据的实验表明,我们的方法优于线性和神经基线,可实现高达12阶的准确HOA RIR估计。对双耳渲染的聆听测试(包括模拟和测量的RIR)进一步证实,所提出的方法比所有基线产生与参考Ambisonics RIR更高的感知相似度。该框架的灵活性和准确性为可扩展声学模拟开辟了新的可能性。
英文摘要
We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spatial capture capabilities, such as irregular or sparse arrays, as classical linear methods fail to reconstruct high-order spatial detail. We introduce a diffusion-based generative framework that models the statistical properties of HOA RIRs. This enables device-agnostic encoding from arbitrary microphone arrays, potentially unseen during data measurement. Our approach incorporates a posterior sampling procedure that enforces consistency between the estimated signals and the measurements while plausibly reconstructing spatial information that is unobservable from the limited measurements alone. Experiments on simulated data demonstrate that our method outperforms linear and neural baselines, achieving accurate HOA RIR estimation up to 12th order. A listening test with binaural renderings, including both simulated and measured RIRs, further confirms that the proposed method yields higher perceptual similarity to reference Ambisonics RIRs than all baselines. The flexibility and accuracy of the proposed framework opens new possibilities for scalable acoustics simulations.
CommentsIWAENC 2026