发表机构
The Australian National University(澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种轻量级参数化FOA编解码器,通过将DOA和扩散度组合为3D指向性向量并采用RVQ联合量化,实现低码率高效编码,在重建质量和下游任务上优于或媲美现有方法。
AI 中文摘要
受沉浸式远程会议和生成式空间音频快速发展的推动,一阶环境立体声(FOA)的高效低码率编码变得越来越重要。在本工作中,我们开发了一种轻量级参数化FOA编解码器,该编解码器保留了标准的定向音频编码分析与合成,同时重新设计了空间元数据量化方案。我们不独立量化到达方向(DOA)和扩散度,而是将它们组合成一个三维指向性向量,并通过残差向量量化(RVQ)跨频带联合量化这些向量。RVQ码本通过分阶段k均值在几分钟内优化,无需反向传播,从而产生恒定比特率(CBR)表示,将元数据速率与频带数量解耦。评估表明,我们的方法在FOA重建方面优于低比特率感知编解码器,在下游声音事件定位与检测任务上与神经编解码器保持竞争力,并且在与不同的外部单声道编解码器配对时保持稳健性能。鉴于其轻量级训练、强大性能和CBR设计,我们认为所提出的方法是未来FOA编解码器研究的一个有利且可复现的基线。
英文摘要
Driven by the rapid growth of immersive teleconferencing and generative spatial audio, efficient low-bitrate coding of first-order Ambisonics (FOA) has become increasingly important. In this work, we develop a lightweight parametric FOA codec that retains the standard Directional Audio Coding analysis and synthesis while redesigning the spatial metadata quantization scheme. Rather than quantizing direction-of-arrival (DOA) and diffuseness independently, we combine them into a 3-D directivity vector and jointly quantize these vectors across frequency bands via residual vector quantization (RVQ). The RVQ codebooks are optimized within minutes via stage-wise k-means without backpropagation, yielding a constant-bitrate (CBR) representation that decouples metadata rate from the number of frequency bands. Evaluations show that our approach outperforms low-bitrate perceptual codecs in FOA reconstruction, remains competitive with neural codecs on the downstream Sound Event Localization and Detection task, and maintains robust performance when paired with different external monaural codecs. Given its lightweight training, strong performance, and CBR design, we consider the proposed method to be a favorable and reproducible baseline for future FOA codec research.