arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Bin2Ambi:从头部追踪双耳音频学习重建Ambisonics声场

Bin2Ambi: Learning Ambisonic Soundfield Reconstruction from Head-Tracked Binaural Audio

Gavin Milner, Nils Peters

arXiv 2609.39732首次发表:更新:

AI 中文总结

本文提出Bin2Ambi任务,利用智能耳机头部追踪数据解决双耳音频方向不确定性,实现向Ambisonics声场的转换,平均方向误差达11.8°,感知质量接近DirAC模型。

AI 中文摘要

用户生成内容已成为最常被消费的内容类型之一。然而,使用消费级硬件捕捉空间音频仍然具有挑战性。鉴于智能耳机的广泛成功,双耳音频有望成为在消费设备上捕捉空间音频的一种有前景的选择,但其固有的信号特性限制了其作为录音格式的可用性。在本文中,我们提出并定义了一个新任务,即双耳到Ambisonics转换(Bin2Ambi)。在我们提出的系统中,我们利用了智能耳机中运动传感器提供的同步捕获的头部追踪数据。我们表明,这些运动数据有助于解决双声道双耳音频因前后定位模糊性和混淆锥中的横向误差而固有的方向不确定性。我们的结果表明,我们的系统学习了方向和扩散场信息,并且头部追踪尤其减少了极端定位误差。客观指标和主观听音测试表明,转换后的Ambisonics声场实现了平均方向误差高达11.8°,且感知空间质量与DirAC真实模型相似。所提出的算法可作为未来改进这一新颖Bin2Ambi任务的基线。

英文摘要

User-generated content has become one of the most-consumed content types. However, capturing spatial audio with consumer hardware is still challenging. Given the widespread success of smart earbuds, binaural audio could be a promising option to capture spatial audio on consumer devices, but its inherent signal characteristic limits its usability as a recording format. In this paper, we propose and define a new task, Binaural to Ambisonics conversion (Bin2Ambi). In our proposed system, we exploit simultaneously captured head-tracking data provided from the motion sensors in smart earbuds. We show that this motion data help resolve the inherent directional uncertainty of two-channel binaural audio due to front-back localization ambiguities and lateral errors in the cone of confusion. Our results show that our system learns directional and diffuse-field information and that head-tracking especially reduces extreme localization errors. Objective metrics and a subjective listening test suggest that the converted Ambisonics soundfield achieves an average directional error of up to $11.8^\circ$ and a perceived spatial quality similar to a DirAC ground-truth model. The proposed algorithm can serve as a baseline for future improvements to this novel Bin2Ambi task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑