arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

听、看与跟踪:面向全模态大模型的时空视听声音事件推理

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang, Yifei Zheng, Minnan Luo

arXiv 2608.09435首次发表:更新:

AI 中文总结

针对现有模型缺失的时空视听推理能力,构建ST-OmniQA基准,提出ST-Omni-R1模型,在该基准上平均语义准确率达77.83%,且空间运动表示可跨基准迁移。

AI 中文摘要

理解动态声源需要联合确定声源是什么、其位置以及随时间的移动方式。然而现有的音频-语言模型常将音频片段表示为全局声学事件,而视觉-语言模型缺乏定位和跟踪单个声源所需的空间音频线索。为评估这一缺失能力,我们引入ST-OmniQA,这是一个时空视听问答基准,由全景视频与同步的一阶Ambisonics(FOA)移动声源音频配对构建而成。该基准包含4万个视频和40万组问答对,分为四个能力层级,涵盖声音事件识别、到达方向、声源距离、运动轨迹以及时间定位的视听推理。基于该基准,我们提出ST-Omni-R1,它整合了FOA衍生的语义与轨迹表示、全景视觉上下文,并通过渐进式课程学习和推理树强化学习进行训练。ST-Omni-R1在四个层级上的平均语义准确率达到77.83%,而评估的最佳基线仅为37.28%。在三个公开空间音频基准上的结果进一步表明,其学习到的空间与运动表示可迁移至ST-OmniQA之外的场景。

英文摘要

Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑