发表机构
Microsoft; Georgia Institute of Technology(微软; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频视觉语言模型评估中时间基准分数衡量问题,提出无标签筛选方法“反转下降”区分模型通道使用情况,如Molmo2和Qwen3-VL,指出综合分数不能反映潜在失败模式,此区分在多基准和任务中成立。
AI 中文摘要
时间视频问答基准上的分数旨在衡量模型是否具有时间理解能力,但它混淆了两个问题。一是任务问题,即问题是否具有时间性,是否需要多个帧及其顺序;二是通道问题,当需要时,模型是从像素中恢复顺序,还是从位置编码(RoPE)中读取顺序。大多数时间分数都无法回答这两个问题,单帧和答案先验往往就能得出分数。该领域的有效性检查、帧打乱敏感性以及从完整视频中获得的准确率,仅涉及任务问题。我们为通道问题贡献了一个无标签筛选方法——反转下降:当视觉序列反转而RoPE保持正向时损失的准确率。它可应用于兼容的时间基准,无需新的注释。配对的反向标签或标签在反转下确定性变换的任务,可区分遵循反向内容的模型和仅因冲突而被打乱的模型。Molmo2从位置读取正向事件的顺序,而Qwen3-VL读取其实际看到的反向事件的视觉顺序。我们称它们为位置主导型和视觉序列主导型。这种区分在两个基准和两个尺度的几个时间任务中都成立,激活修补表明这是一种真实的内部属性,而非冲突的假象。这种区分很重要,因为两个通道在相反的输入上会失败,所以两个分数相似的模型不可互换,即综合分数不能反映潜在的失败模式。
英文摘要
A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1. The task question: is the question even temporal, does it need several frames and their order? and 2. The channel question, when it does, does the model recover the order from the pixels, or read it off the positional encoding (RoPE)? Most of a temporal score answers neither, a single frame and answer priors often carry it. The field's validity checks, frame-shuffle sensitivity and the accuracy gained from the full video, speak only to the task question. We contribute a label-free screen for the channel question, the reversal-drop: the accuracy lost when the visual sequence is reversed while RoPE remains forward. It can be applied to compatible temporal benchmarks without new annotations. Paired reverse labels, or tasks whose labels transform deterministically under reversal, distinguish models that follow reversed content from those merely disrupted by the conflict. Molmo2 answers the forward event reading order off positions, while Qwen3-VL answers the reversed event it actually sees, reading visual order (comparatively). We call them position-dominant and visual-sequence-dominant. The split holds across two benchmarks and several temporal tasks at two scales, and activation patching shows it is a real internal property, not an artifact of the conflict. The distinction matters, the two channels fail on opposite inputs so two models with similar score are not interchangable, i.e. an aggregate score does not reflect potential failure modes.
Comments9 pages, 11 pages supplemental