发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出 VIDEOHARNESS-RSI,通过递归搜索可执行上下文构建程序,在冻结 VLM 上实现长视频理解的改进,且所选 harness 可迁移至其他长视频基准。
AI 中文摘要
长视频理解高度依赖如何从更长的视频中构建有限的模型上下文。现有方法通过压缩、检索、记忆和智能体式证据获取来改进这一过程,但这些机制通常作为手动设计的推理系统的一部分引入,或与其他组件共同优化,这使得难以分离一个更简单的问题:仅通过改进可执行的上下文构建程序能获得多少收益?我们通过 VIDEOHARNESS-RSI 研究该问题,这是围绕冻结视觉语言模型(VLM)递归搜索可执行上下文构建器的受控基线。外层提议器利用先前的程序、评估结果和执行轨迹生成候选 harness,这些候选 harness 经端到端执行和评估后,成功的变体会被保留以进行进一步搜索。这使得长视频理解成为自动化 harness 设计的受控实例:可搜索对象是可执行程序结构,而回答模型和接口保持固定。从均匀采样开始,递归 harness 搜索持续发现改进空间,并优于多个较弱的手动设计基线;从更强的手动设计基线开始,相同的 RSI 过程可进一步提升性能。所选 harness 还可迁移到其他长视频基准,无需进一步搜索。这些结果共同确立了可执行上下文构建作为一个独特的优化层,并为研究冻结 VLM 周围的 harness 发现和迁移提供了可复现的基线。
英文摘要
Long-video understanding depends not only on the capability of a vision-language model (VLM), but also on how its limited context is constructed from a much longer video. Existing systems typically introduce hand-designed sampling, retrieval, memory, or agentic control strategies, making the context-construction program itself difficult to study as an independent optimization target. We introduce VideoHarness-RSI, a controlled framework that recursively searches executable context constructors around a frozen VLM while keeping the answering model and interface fixed. We study this baseline under complementary weak- and strong-initialization regimes. From a weak uniform constructor, recursive search progressively discovers more structured context-construction programs; from a stronger AKS harness, the same process further advances an already competitive hand-crafted frontier. The resulting harness retains its advantage under a matched cumulative visual-token control and transfers directly to additional long-video benchmarks without further search. Together, these results establish executable context construction as a distinct optimization layer and provide an auditable baseline for studying harness discovery, transfer, and efficiency around frozen VLMs.