发表机构
Ben-Gurion University of the Negev(内盖夫本-古里安大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明冻结视频语言模型已隐式编码证据就绪性信号,可线性解码并用于答案时序门控,在匹配时长下提升准确率最多9.75个百分点。
AI 中文摘要
流式视频-语言模型不仅需要决定回答什么,还需要决定当前问题所需的证据是否已经到达。现有系统将该决策学习为一个单独的触发器;我们探究一个未经修改的模型是否已经计算了该决策。我们证明,冻结的VideoLLMs携带一个线性可读的证据就绪性信号,该信号基于带时间戳的证据而非模型输出进行标注。在共享字节相同评估的七个模型中,该信号均可解码(在最严格的未就绪采样下,AUROC为0.733-0.905,而此时拟合的时钟接近随机水平),并且在不包含某个基准家族任何视频片段的情况下拟合的探针仍能读取该家族。该信号是问题条件化的:在字节相同的窗口上,仅改变问题即可在66.1%的配对中反转读出结果,而所有问题无关的对照在构造上均处于随机水平。模型可能回答错误但仍编码就绪性:在错误答案中AUROC仍为0.722。就绪性在延迟匹配的答案选择上也优于不确定性估计器及其监督组合,并且比置信度更紧密地跟踪独立的人类判断。已发布的流式触发器也是线性读出,但在其自身基础模型的激活上训练的触发器与就绪性近似正交,并且解码就绪性的准确性远低于探针。我们将读出转化为就绪性门控,这是一种答案时序策略,在匹配的视频时长下将准确率提高最多+9.75个百分点,且计算开销可忽略不计。其收益大小随任务提供的准确率余量而变化:在26种配置中,收益跟踪该余量,并且一种在相同像素上移动该余量的干预也会随之移动收益。
英文摘要
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.