发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对 VideoLMs 时间理解中监督与视频表示不匹配的问题,提出 VT-Contrast 表示级时间反事实目标,通过对比视频 token 的顺序视图与反事实提升时间理解性能,且无需改动架构。
AI 中文摘要
按顺序观看帧并不意味着能表示时间。现代 VideoLMs 接收有序视频流,但它们的主要监督作用于生成的文本,而非应首先呈现事件动态的视频 token 表示。这种不匹配使模型能从物体、场景和语言先验等捷径中学习时间答案,而无需内部视频表示捕捉事件进展。为解决此问题,我们提出 VT-Contrast,一种针对 VideoLMs 的表示级时间反事实目标。其设计明确时间监督应作用于何处以及应揭示何种时间差异:VT-Contrast 监督选定的后期层最后一帧视频 token(这些 token 应在语言生成前整合时间信息),并对比顺序保留视图与按 Kendall tau 距离分级的同视频重排反事实。该方法无需架构改动,可兼容多种 VideoLM 训练任务,并在时间理解基准上提升了整体性能,代码可在指定 URL 获取。
英文摘要
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This mismatch allows models to learn temporal answers from shortcuts such as objects, scenes, and language priors, without requiring internal video representations to capture event progression. To address this, we propose VT-Contrast, a representation-level temporal counterfactual objective for VideoLMs. Its design asks where temporal supervision should act and what temporal differences it should expose. VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance. It requires no architectural changes, is compatible with diverse VideoLM training tasks, and improves overall performance across temporal understanding benchmarks. Our code is available at https://github.com/ANDgate99/VT-Contrast.
CommentsAccepted to EMNLP 2026 (main)