arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08542eess.AS

基于相对房间冲激响应估计的空间音频编码

Spatial Audio Coding Through Relative Room Impulse Response Estimation

  • Orange Research(Orange研究)
  • IRISA

机构由 AI 辅助整理,请以论文原文为准。

Nour Bouayed, Adrien Llave, Jérôme Daniel, Pascal Scalart

中文总结 AI 辅助

本文提出一种基于相对空间房间冲激响应盲估计的HOA编码方案,利用其稀疏性实现高效参数化表示,在单通道下压缩率优于IVAS且质量相当或更优。

中文摘要 AI 辅助

沉浸式虚拟聆听依赖于高阶环境立体声(HOA)等空间音频技术,这些技术将声音场景表示为多通道信号。随着所需空间分辨率的提高,通道数量也随之增加,这使得高效压缩对于在带宽受限网络上传输至关重要。此外,为了便于网络运营商部署,沉浸式音频编码的目标比特率理想情况下应保持在当前分配给VoLTE音频服务的25 kbps附近。最先进的参数化编解码器,例如最近标准化的沉浸式语音和音频服务(IVAS)编解码器,通过传输空间元数据以及减少数量的传输通道来实现压缩。然而,最近的研究表明,IVAS在混响内容上的性能会下降,尤其是在低比特率下,这一局限性表明其无法准确建模房间声学。在本文中,我们提出了一种新颖的HOA编码方案,该方案基于相对空间房间冲激响应(ReSRIR)的显式和盲估计,使用HOA信号的波束成形版本作为参考信号。通过利用估计的ReSRIR的结构和稀疏性,我们推导出一种用于沉浸式音频编码的高效参数化表示。实验评估表明,所提出的方法在单传输通道情况下比IVAS实现了更高的压缩率,同时保持了相当或略好的质量。

英文摘要

Immersive virtual listening relies on spatial audio technologies such as Higher-Order Ambisonics (HOA), which represent sound scenes as multichannel signals. As the desired spatial resolution increases, so does the number of channels, making efficient compression essential for transmission over bandwidth-limited networks. Moreover, to facilitate deployment by network operators, the target bitrate for immersive audio coding should ideally remain close to the 25 kbps currently allocated to VoLTE audio services. State-of-the-art parametric codecs, such as the recently standardized Immersive Voice and Audio Services (IVAS) codec, achieve compression by transmitting spatial metadata together with a reduced number of transport channels. However, recent studies have shown that IVAS performance degrades on reverberant content, particularly at low bitrates, a limitation that suggests its inability to accurately model room acoustics. In this paper, we propose a novel HOA coding scheme based on the explicit and blind estimation of the Relative Spatial Room Impulse Response (ReSRIR), using a beamformed version of the HOA signal as a reference signal. By exploiting the structure and sparsity of the estimated ReSRIR, we derive an efficient parametric representation for immersive audio coding. Experimental evaluations show that the proposed method achieves higher compression than IVAS in the single-transport-channel regime, while maintaining comparable to slightly better quality.

补充信息

↑