arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

这里的世界是立体的:学习动态空间对应关系以进行沉浸式联合视频-音频生成

Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Hanmo Chen, Chengcheng Liu, Tianxiao Chen, Zheyu Zhang, Siming Zheng, Jinwei Chen, Xu Yang, Cheng Deng, Bo Li, Peng-tao Jiang

arXiv 2609.38748首次发表:更新:

发表机构

Xidian University; vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.(西安电子科技大学; vivo蓝心影像实验室,vivo移动通信有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有联合视频-音频生成模型缺乏立体声沉浸感的问题,提出StereoBind框架,通过三种机制绑定视觉运动与立体声生成,并构建数据集和基准,实验证明其显著提升空间对齐且保持整体质量。

AI 中文摘要

最近的联合视频-音频生成模型已经实现了强大的语义对应和时间同步。然而,AR/VR和交互式游戏等应用进一步需要立体声以提供沉浸感,这一点在很大程度上仍被忽视。有效的立体声要求感知到的声音位置与对应视觉源的运动保持一致。我们将这一特性称为动态空间对应关系,并提出StereoBind,一个将视觉源运动绑定到立体声生成的框架。StereoBind通过三种互补机制使用运动轨迹来协调视觉运动和立体声音频。视觉运动绑定建立了基于源感知的视听对应关系,空间轨迹编码器捕获绝对源位置,残差轨迹RoPE对相对运动进行建模。为了监督和评估,我们构建了StereoWorld-29K,一个带有配对运动轨迹的大规模立体声音频-视频数据集,以及StereoWorldBench用于衡量视听空间一致性。实验表明,与现有模型相比,StereoBind在立体声音频生成中显著提高了空间对齐,同时保持了整体视听质量。

英文摘要

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑