发表机构
Beijing University of Chemical Technology; Baidu, Inc.(北京化工大学; 百度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长视频智能体现有统一证据获取方式的缺陷,提出无需训练的VESTA智能体,通过策略引导多策略检索等机制,在多个长视频基准上取得了显著性能提升。
AI 中文摘要
现有的长视频智能体通过统一行为获取证据,未考虑所需证据是否集中、是否需要广泛覆盖或是否必须区分竞争假设,这可能导致在实质性推理开始前就失败。为每个问题规定精细的解决方案并非令人满意的补救措施,因为它限制了自主探索。我们提出VESTA,这是一种无需训练的长视频智能体,组织为路径条件的获取-验证-整合循环。在探索前,意图路由器会推断证据获取策略(针对共享视觉-语音场景索引的聚焦、召回或对比检索),以及配置探索期间维护的证据视图的证据核算策略。策略引导检索产生临时参考,多模态证据操作将其转换为观测结果,同时推理器可自由验证这些参考、使用中间发现重新查询或检查检索集之外的区域。时间证据账本将观测结果整合为时间位置、来源、覆盖范围、冲突、验证结果和假设支持的自适应压缩视图,暴露缺失和未解决的证据以指导后续获取;最终确定优先考虑已验证的观测结果。在Video-MME-v2上,VESTA的平均准确率比VideoARM提高2.7个百分点,且在所有6个报告指标上均有提升;在共享查询时间模型下,其在LongVideoBench长子集上提高6.9个百分点,在LVBench上提高1.5个百分点,在EgoSchema上与VideoARM持平。
英文摘要
Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for every question is not a satisfactory remedy, as it restricts autonomous exploration. We propose VESTA, a training-free long-video agent organized as a route-conditioned acquire--verify--consolidate loop. Before exploration, an intent router infers an evidence-acquisition policy---focused, recall, or contrastive retrieval over a shared visual--speech scene index---together with an evidence-accounting policy that configures the evidence view maintained during exploration. Policy-steered retrieval yields provisional references that multimodal evidence operations convert into observations, while the Reasoner remains free to verify them, re-query using intermediate findings, or inspect regions outside the retrieved set. A temporal evidence ledger consolidates observations into an adaptive, compressed view of temporal location, provenance, coverage, conflicts, verification outcomes, and hypothesis support, exposing missing and unresolved evidence to guide subsequent acquisition; finalization prioritizes verified observations. On Video-MME-v2, VESTA improves average accuracy by 2.7 points over VideoARM and gains across all six reported metrics. On LongVideoBench, EgoSchema, and LVBench under shared query-time models, it improves by 6.9 points on the LongVideoBench long subset and 1.5 on LVBench, and matches VideoARM on EgoSchema.
Comments15 pages, 5 figures, 6 tables (main paper with appendix)