生成式学习用于Ambisonic上混
Generative Learning for Ambisonic Upscaling
查看机构详情
- ECE School Ben-Gurion University of the Negev(内盖夫本-古里安大学电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究将Ambisonic上混视为生成任务,采用流匹配和基于分数的生成模型恢复混响语音的空间信息,实验表明流匹配在混响场景中性能最优。
中文摘要 AI 辅助
Ambisonic上混(AU)旨在通过从低阶观测中估计高阶Ambisonic(HOA)分量来增强声场的空间分辨率。尽管深度学习和基于模型的策略已被考虑用于AU,但这两种方法在现实场景中均表现出显著的性能下降,因为混响声场违反了判别性映射固有的方向稀疏性。在这项工作中,我们将AU视为一个生成任务而非确定性重建,扩展生成建模以专门针对混响语音中空间信息的恢复。我们研究了两种主流的连续时间生成范式,将基于分数的生成模型和流匹配适配到这些复杂的声学环境中。我们提供了广泛的数值研究,在各种声学场景中将我们的方法与最先进的基线进行比较。此外,我们进行了主观听力测试,以评估所提出的生成框架在各种混响场景中的感知质量和空间准确性。研究表明,流匹配在所有混响设置中始终优于其判别性对应方法和基于扩散的范式。
英文摘要
Ambisonics Upscaling (AU) aims to enhance the spatial resolution of sound fields by estimating high-order Ambisonics (HOA) components from low-order observations. While deep learning and model-based strategies have been considered for AU, both approaches exhibit significant performance degradation in realistic scenarios, where reverberant sound fields violate the directional sparsity inherent to discriminative mappings. In this work, we address AU as a generative task rather than a deterministic reconstruction, expanding generative modeling to specifically target the recovery of spatial information in reverberant speech. We investigate two dominant continuous-time generative paradigms, adapting both Score-based Generative Model and Flow Matching to these complex acoustic settings. We provide an extensive numerical study comparing our methods against state-of-the-art baselines in various acoustic scenarios. Additionally, we conduct subjective listening tests to evaluate the perceived quality and spatial accuracy of the proposed generative framework across various reverberant scenarios. The studies reveal that Flow Matching consistently outperforms both its discriminative counterparts and Diffusion-based paradigms in all reverberant settings.