发表机构
The Hong Kong University of Science and Technology; Noiz AI; MetaX; Shanghai Jiao Tong University(香港科技大学; Noiz AI; MetaX; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WorldSonus提出交互式视频到音频框架,通过流式因果自回归扩散架构实现低RTF实时生成,支持交互控制与空间对齐,在开放域基准上达到或超越最先进模型。
AI 中文摘要
世界模型的最新进展已能够实现越来越逼真的视觉合成。然而,这些生成的环境在很大程度上仍然是无声的。为世界模型赋予声音面临三个核心挑战:实时生成以跟上交互式视频流的节奏、交互式控制以响应生成过程中的声音指令、以及空间对齐的立体声以反映场景几何和相机运动。为了应对这些需求,我们提出了WorldSonus,一个专为世界模型中的实时空间声音合成而设计的交互式视频到音频框架。在实时生成方面,WorldSonus采用流式因果自回归扩散架构,以0.41的低实时因子(RTF)合成音频块。在交互控制方面,我们整合了一个以音频为中心的描述流水线,并采用分块索引的提示调度,从而在生成过程中实现对声音事件的动态操控。在空间对齐方面,我们利用从多样化的立体声和环境声数据中精心策划的高质量立体声监督。大量实验表明,尽管WorldSonus是为世界模型量身定制的,但它能有效地泛化到开放域的视频到音频基准测试中,在声学质量和空间对齐方面均达到或超越了最先进的双向模型。项目页面:此https URL。
英文摘要
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/
Comments25 pages, 4 figures, 16 tables. Project page: https://noizai.github.io/WorldSonus/