AudioWorldSim:用于世界模型的真实双耳音频数据集
AudioWorldSim: Realistic Binaural Audio Datasets For World Models
浏览论文内容
中文总结 AI 辅助
本研究提出开源平台AudioWorldSim,作为SoundSpaces 2.0的自定义扩展,用于生成真实双耳音频数据集,助力基于音频的机器学习尤其是世界模型研究,相关资源已公开以提升可重复性。
中文摘要 AI 辅助
本技术报告介绍AudioWorldSim,这是一个开源平台,旨在生成真实的双耳音频数据集,推进基于音频的机器学习研究,尤其是世界模型领域。AudioWorldSim作为Meta的SoundSpaces 2.0平台的自定义扩展构建,利用其全面的声学框架,专注于随机智能体导航的自动展开,还对连续声音的组合方式进行了关键修复。AudioWorldSim已向研究界公开提供,可通过指定网址获取,以促进研究可重复性。
英文摘要
This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a custom extension of Meta's SoundSpaces 2.0 platform, AudioWorldSim leverages their comprehensive acoustics framework, but focuses on the automatic rollout of random agent navigations, as well as implements crucial fixes to how continuous sound is composed. AudioWorldSim is made publicly available to the research community at https://github.com/Luizerko/AudioWorldSim to facilitate reproducibility.
发表机构
- VISGRAF
- IMPA(巴西纯数学与应用数学国家研究所)
机构由 AI 辅助整理,请以论文原文为准。