arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36295cs.SDcs.MMeess.AS

将任意视频转化为沉浸式视听体验

Enabling Immersive Audio-Visual Experience from Any Video

Zitong Lan, Mutian Tong, Jiatao Gu, Mingmin Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

OmniDream是一个无需训练的框架,通过对象为中心的音频表示将无声单目视频转化为沉浸式视听体验,实现空间对齐的音频与视觉场景,提升音视频对齐和沉浸感。

中文摘要 AI 辅助

大多数视频仅捕捉狭窄的视野且不提供空间音频,限制了它们所能提供的沉浸感。最近的视频生成模型可以将透视视频扩展为全景视频,但并未提供相应的空间声景。没有空间一致的音频,这些扩展的视觉世界仍然是不完整的。本文提出了OmniDream,一个无需训练的框架,可将无声的单目视频转化为沉浸式视听体验,其中观众可以自由环顾四周,而声音则与视觉场景保持空间对齐。OmniDream的核心是一种以对象为中心的音频表示,它将每个声源的固有音频内容与其场景相关的声学效果分离,从而实现独立的音频生成、基于物理的传播效果模拟以及灵活的空间音频渲染。实验表明,与基线相比,该方法在音视频对齐、空间正确性和感知沉浸感方面均有提升。示例可在以下网址获取:https URL

英文摘要

Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑