arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34363cs.CVcs.AIcs.MM

SyncRA:在全模态模型中学习时间对应关系

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • University of Alberta(阿尔伯塔大学)
  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Zelong Xu, Yan Li, Wenhe Hu, Xiyang Hu

AI总结:

针对全模态模型难以关联同刻音视频的问题,提出轻量级方法SyncRA,通过对比中间表示对齐匹配时刻并分离不匹配时刻,无需额外标注,在多个模型和基准上一致优于仅答案微调,显著提升时间对应与问答性能。

AI中文摘要:

近期的全模态模型展现出对音频和视觉输入的强大感知能力,但往往难以将同一时刻所听到的内容与所看到的内容联系起来。这种时间对应关系的薄弱可能导致模型将语音线索与错误的视觉场景关联起来,从而基于错误的音视频配对产生看似合理的答案。我们通过受控的时间交换实验诊断了这一问题,结果表明模型的答案并不能可靠地跟随这些配对的变化。为解决此问题,我们提出了同步引导表示对齐(SyncRA),一种用于增强音频与视觉之间时间对应关系的轻量级方法。具体而言,SyncRA 在每个视频内部对比中间音视频表示,对齐匹配的时刻同时分离不匹配的时刻,从而在共享的全局上下文中捕获局部时间对应关系。该目标直接从现有输入时序中获取监督信号,无需额外标注,且推理过程保持不变。我们在四个覆盖不同规模和架构的开源全模态模型上,对五个公开视频基准进行了 SyncRA 评估。SyncRA 在所有模型-基准组合上均一致优于仅答案微调方法,同时在受控评估中显著提升了追踪变化音视频配对的能力。这些结果表明,轻量级、有针对性的监督能够有效增强时间对应关系,并转化为音视频问答任务的广泛性能提升。

英文摘要:

Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.

补充信息

↑