arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06837cs.SDcs.LG

BinauralVAE:用于世界模型的空间音频重建

BinauralVAE: Spatial Audio Reconstruction For World Models

Luis Vitor Zerkowski, Luiz Velho

首次发表
浏览论文内容

中文总结 AI 辅助

BinauralVAE提出一个开源流水线,利用多种变分自编码器架构重建空间音频,为基于音频的世界模型奠定状态表示基础,以补充视觉感知。

中文摘要 AI 辅助

具身人工智能历来在很大程度上依赖于视觉感知,导致多种以视觉为中心的世界模型激增。然而,这种依赖未能完整捕捉空间理解,并且在存在视觉遮挡、低光照条件或断电等场景的环境中甚至可能暴露出脆弱性——在这些场景中,声学信息成为空间感知和导航的关键替代方案。尽管具有潜力,但对现实空间音频的研究,特别是以音频为中心的世界模型的开发,仍然稀少。在本技术报告中,我们介绍了BinauralVAE:一个灵活、开源的流水线(此https链接),它探索了多种用于空间化音频重建的模型,从基础基线发展到先进的、数学基础扎实的架构。我们的方法评估了多种变分自编码器架构——包括复数值变体——以学习双耳信号的鲁棒潜在表示。与AudioWorldSim一同开发,我们的方法利用模拟机器人在环境中导航时捕获的现实声学数据。该流水线为未来基于音频的世界模型中的状态表示奠定了基础,旨在映射导航动作与其产生的声学后果之间的直接因果关系,并有助于使声音成为空间知识获取的基本补充模态。

英文摘要

Embodied artificial intelligence has historically very much relied on visual perception, leading to a proliferation of multiple vision-centric world models. However, this reliance fails to capture spatial understanding in its entirety and can even present vulnerabilities in environments with visual occlusions, low-light conditions, or blackouts-scenarios, where acoustic information becomes a critical alternative for spatial awareness and navigation. Despite its potential, research into realistic spatial audio and particularly the development of audio-centric world models remains sparse. In this technical report, we introduce BinauralVAE: a flexible, open-source pipeline (https://github.com/Luizerko/BinauralVAE) that explores multiple models for spatialized audio reconstruction, progressing from fundamental baselines to advanced, mathematically grounded architectures. Our approach evaluates various Variational Autoencoder architectures -- including complex-valued variants -- to learn robust latent representations of binaural signals. Developed alongside AudioWorldSim, our methodology leverages realistic acoustic data captured as a simulated robot navigates an environment. This pipeline establishes a foundation for state representation in a future audio-based world model, designed to map the direct causal connection between navigational actions and their resulting acoustic consequences, and helping to enable sound as an essential complementary modality for spatial knowledge acquisition.

发表机构

  • VISGRAF
  • IMPA(巴西纯数学与应用数学国家研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑