arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08315cs.CVcs.MM

你的视觉语言模型(VLM)已经知道时间:通过提问“是/否”实现无需训练的时间定位

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No

Ji Huang, Barry Devereux, Hui Wang

AI总结:

针对VLM时间定位任务的自信错误问题,提出无需训练的FV-Action方法,通过从粗到细的二元问题扫描替代时间戳回归,在多个基准数据集上大幅提升了时间定位性能。

AI中文摘要:

能够可靠识别事件的多模态大语言模型(LLM)仍无法准确说出事件发生的时间。当要求输出时间戳时,性能强劲的视觉语言模型(VLM)在Charades-STA数据集上的R@0.5仅达到3.8%,且其77%至80%的错误预测具有低输出熵:模型是自信地出错,而基于熵的错误检测效果低于随机分类器。我们证明,这种失败源于任务接口,而非感知能力。在保持模型权重固定的情况下,将时间戳回归替换为从粗到细的二元问题扫描,仅将第一个标记的概率用作排序依据,可在四个冻结的骨干网络上将R@0.5提升28至50个百分点。剩余的失败可分解为两个可测量的维度:随骨干网络变化的感知维度,以及可通过输出窗口与事件宽度的比率解析预测的几何维度。基于此分析构建的无需训练的方法FV-Action,在Charades-STA数据集上达到56.8%的R@0.5,优于同骨干网络的原生定位管道及该基准上最强的无需训练结果;它在TACoS上零样本评估时超越所有经TVG训练的模型,且在ActivityNet Captions和QVHighlights上较直接预测有所提升,全程无任何时间监督。

英文摘要:

Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as $3.8\%$ R@0.5 on Charades-STA, and $77$ to $80\%$ of their wrong predictions carry low output entropy: the models are confidently wrong, and entropy-based error detection stays below a random classifier. We show that this failure lives in the task interface, not in perception. Holding the weights fixed, replacing timestamp regression with a coarse-to-fine scan of binary questions, whose first-token probabilities are consumed only as a ranking, raises R@0.5 by $28$ to $50$ points across four frozen backbones. The residual failures decompose into two measurable axes: a perception axis that moves with the backbone, and a geometry axis that is analytically predictable from the ratio of the output-window and event widths. FV-Action, the training-free method built on this analysis, reaches $56.8\%$ R@0.5 on Charades-STA, above the same backbone's native grounding pipeline and the strongest training-free result on this benchmark; it surpasses every TVG-trained model evaluated zero-shot on TACoS, and improves over direct prediction on ActivityNet Captions and QVHighlights, with no temporal supervision at any stage.

↑