发表机构
NAAMII, Nepal; University College London, UK; University of Aberdeen, UK(NAAMII(尼泊尔); 伦敦大学学院; 阿伯丁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FORESIGHT双流架构,利用流式VLM的预测能力,无需训练即可动态规划未来计算,在多个基准上显著提升性能。
AI 中文摘要
现有的流式视觉语言模型(VLMs)持续感知并推理视觉流,但其计算路径在推理过程中保持固定。因此,它们无法根据不断变化的场景动态调整计算,而不同的未来事件需要不同级别和形式的感知。我们证明,流式视觉语言模型天生具备预测近期未来的能力,并利用这一能力以无需训练的方式动态配置未来计算。然而,实现这种预测性计算极具挑战性:未来预测必须足够可靠以指导计算,规划必须与流式推理并行运行,且在线重配置的开销必须可忽略不计。为应对这些挑战,我们提出了FORESIGHT,一种双流架构,包含两个共享权重、输入编码器和KV缓存的孪生(Siamese)LLM。第一个LLM持续处理传入的令牌,而第二个LLM则领先于流运行,以预测未来上下文、规划未来计算并生成任务响应,且不中断流式推理。每个规划决定何时进行下一次推理、届时检查什么内容以及采样的密度,将瞬时证据与持久控制分离。所得计算规划通过一种高效的在线重配置协议执行,该协议采用模式引导解码和轻量级基于差异的更新,从而实现低开销的动态适应。在冻结的Qwen3-VL-8B骨干网络上,FORESIGHT在OmniPro在线评估中实现了23.0的平均联合F1分数,比最强训练基线高出9.5%,同时在StreamingBench上比骨干网络提升6.7,在OVO-Bench上提升15.4,当证据在视频流中较晚出现时,最大提升达到18.7。我们的源代码将公开发布。
英文摘要
Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.