AVTrace:诊断全模态模型中的音视频时间推理
AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models
浏览论文内容
中文总结 AI 辅助
AVTrace是一个音视频时间推理诊断套件,通过评估五个全模态模型,发现其在同步性验证上均低于基线,并揭示了时间后训练能提升部分模型性能,强调语义重叠不能替代时间定位。
中文摘要 AI 辅助
全模态模型能够描述视频内容,但它们能否在时间上定位事件、保持事件顺序并判断音视频同步性?我们引入了AVTrace(音视频时间推理评估与能力评测),这是一个银标准诊断套件,涵盖起始与跨度定位、同步性、下一步预测、跨模态定位、链式解析以及事件条件理解。它包含34,114个训练示例,以及类别平衡的开发集和测试集,分别包含3,500和7,000个示例。我们在各自输入配置下,使用参考盲响应归一化后进行确定性评分,评估了五个开放全模态模型。所有五个现成系统在同步性验证上的得分均低于测试集多数标签基线0.556,并在链式解析、事件条件定位和理解上得分较低。开发集扰动揭示了Qwen3-Omni-30B对模态移除和视觉输入处理变化的任务依赖性敏感性,但未隔离其根本原因。参数高效的时间后训练在多个基准指标上提升了Gemma4-E4B-it的性能。在三个外部图像基准上,任务指标变化不大,包括一些退化,而教师强制困惑度下降。综合这些发现表明,语义参考文本重叠不应被视为时间定位的代理指标,且AVTrace能够识别任务特定弱点,同时为时间后训练提供测试平台。
英文摘要
Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split's majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.
发表机构
- Institute of Advanced Intelligence and Computing (IAIC), A*STAR(A*STAR 高级智能与计算研究所)
- Centre for Frontier AI Research (CFAR), A*STAR(A*STAR 前沿人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。