面向野外360°语音场景的视觉引导空间音频生成
Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes
AI总结:
针对野外360°语音场景空间音频采集质量受限问题,本文提出视觉引导的FOA语音空间化方法,构建YT-SPEECH数据集,采用Localizer-Renderer框架实现定向FOA信号重建,提升了音频相关性能。
AI中文摘要:
空间音频是沉浸式360°媒体的核心组成部分,但在现实世界以语音为主的场景中,高质量的空间采集仍存在局限。本文研究野外环境下视觉引导的一阶Ambisonics(FOA)语音空间化任务:给定对齐的360°视频和全向音频轨道,我们恢复缺失的定向FOA分量。为支持该任务,我们引入YT-SPEECH,这是一个从YouTube整理的面向语音的360°视频-FOA数据集。我们提出两阶段的Localizer-Renderer框架,其中一个视听分割骨干网络提供逐帧空间热图,条件复域U-Net从全向通道重建定向FOA信号;基于置信度的门控策略可在模糊声学条件下稳定条件输入。实验表明,与消融变体及现有方法相比,该方法在重建保真度、空间准确性和感知语音质量方面均有提升。
英文摘要:
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented $360^\circ$ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.