HelixWorld:一种实时交互式视听世界模型
HelixWorld: A Real-time Interactive Audio-Visual World Model
浏览论文内容
中文总结 AI 辅助
HelixWorld提出首个实时交互视听世界模型,通过双向教师蒸馏实现24 FPS无漂移的视听联合推演,并引入HelixBench评估空间声学一致性,在保持视觉性能的同时显著提升声学沉浸感。
中文摘要 AI 辅助
世界模拟本质上是多感官的,要求实时同步的视觉和声学动态。然而,当前主流的交互式世界模型仍然严格保持沉默,仅专注于视觉渲染和控制,而忽略了声学维度。我们提出了HelixWorld,一种实时交互式视听世界模型,其中视觉场景和基于摄像机的空间立体声在用户交互下原生地共同演化。我们构建了一个具有真实立体声学和度量相机位姿的高保真空间视听数据集,并在此基础上预训练了一个以6自由度相机轨迹和用户动作作为条件的双向教师模型。为了实现低延迟的因果交互,我们通过在线轨迹蒸馏损失将教师模型蒸馏为少步流式学生模型,在单个GPU上以24 FPS的速度维持无漂移的联合视听推演。此外,我们形式化了空间声学一致性,并引入了HelixBench来评估合成声场是否忠实地跟踪动态视点运动。大量实验表明,HelixWorld在视觉保真度和响应性方面与最先进的无声世界模型相匹配,同时在相机对齐的空间声学沉浸感方面显著超越了现有基线。
英文摘要
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Noiz AI
机构由 AI 辅助整理,请以论文原文为准。