arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于扩散后验采样的阵列无关 Ambisonics 编码

Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling

Amit Milstein, Nir Shlezinger, Boaz Rafaely

arXiv 2608.24558首次发表:更新:

AI 中文总结

针对现有 Ambisonics 编码方案受固定阵列限制的问题,提出生成框架 ADEPS,将物理采集模型嵌入推理过程,可补偿阵列失真并实现任意阵列零样本编码,性能优于传统基线。

AI 中文摘要

空间音频通过重现三维声场提升用户沉浸感,Ambisonics 是广泛采用的表示方法。尽管 Ambisonics 理论上独立于录音设备,但实际麦克风阵列会引入依赖硬件的编码伪影,现有数据驱动方案缺乏灵活性,通常受限于固定阵列几何结构。为克服这些局限,我们提出 ADEPS,一种将物理采集模型显式嵌入推理过程的生成框架。利用该公式,ADEPS 可有效补偿阵列特定失真,同时支持任意阵列拓扑的零样本编码。我们仅以目标 Ambisonic 表示无监督训练底层生成先验,在多样模拟和真实麦克风阵列上的大量评估表明,ADEPS 在空间保真度和频谱质量上始终优于传统线性和参数基线。

英文摘要

Spatial audio enhances user immersion by reproducing 3D sound fields, with Ambisonics being a widely adopted representation. While Ambisonics is theoretically independent of the recording setup, practical microphone arrays introduce hardware-dependent encoding artifacts. Moreover, existing data-driven solutions lack flexibility, as they are typically restricted to fixed array geometries. To overcome these limitations, we propose ADEPS, a generative framework that explicitly embeds the physical acquisition model into the inference process. By leveraging this formulation, ADEPS effectively compensates for array-specific distortions while enabling zero-shot encoding across arbitrary array topologies. We train the underlying generative prior in an unsupervised manner solely on target Ambisonic representations. Extensive evaluations across diverse simulated and real microphone arrays demonstrate that ADEPS consistently outperforms both traditional linear and parametric baselines in spatial fidelity and spectral quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑