发表机构
The Hong Kong University of Science and Technology; Baidu Inc.; Alibaba Group; Xidian University(香港科技大学; 百度公司; 阿里巴巴集团; 西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对长视频理解的帧选择局限,提出GenEvA框架,通过查询条件化分布聚合跨帧潜在证据,在四个基准和两个视频多模态大语言模型主干上显著提升性能且开销极低。
AI 中文摘要
长视频理解通常会将视频压缩为少量帧或视觉标记以生成答案。现有紧凑流水线聚焦于保留相关视觉内容作为显式证据,但证据的可用性并不意味着能整合不同时刻的互补线索来回答问题。本文的核心思路是在生成前将选定帧组织为与查询相关的跨帧证据,我们将这一后选择阶段形式化为潜在证据接口,并实例化为GenEvA(Generative Latent Evidence Aggregation,生成式潜在证据聚合),一种分布引导的潜在证据聚合框架。具体而言,GenEvA使用查询条件化的证据分布将聚合聚焦于相关帧,从各帧的特定信息中形成紧凑的跨帧潜在证据;由于跨帧整合并非始终必要,同一分布会决定是否插入这一潜在补充。在四个基准和两个视频多模态大语言模型(Video-MLLM)主干上,GenEvA始终优于匹配帧的基线方法:在8帧设置下,它将四个基准的LLaVA-Video平均性能提升5.2个百分点,Qwen2.5-VL在LVBench上的准确率提升10.1个百分点;这些增益仅需0.11%至0.40%的平均视频标记开销,进一步分析显示其具有任务感知分配特性,并受益于自适应证据调用(Adaptive Evidence Invocation)。
英文摘要
Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet making evidence available does not ensure that complementary cues across moments are integrated for answering. Our key idea is to organize selected frames into query-relevant cross-frame evidence before generation. We formulate this post-selection stage as a latent evidence interface and instantiate it with GenEvA ($\textbf{Gen}erative$ $Latent$ $\textbf{Ev}idence$ $\textbf{A}ggregation$), a distribution-guided latent evidence aggregation framework. Specifically, GenEvA uses a query-conditioned evidence distribution to focus aggregation on relevant frames, forming compact cross-frame latent evidence from their frame-specific information. Since cross-frame integration is not always needed, the same distribution determines whether to insert this latent complement. Across four benchmarks and two Video-MLLM backbones, GenEvA consistently improves matched-frame baselines. At 8 frames, it raises the four-benchmark LLaVA-Video average by $+5.2$ points and Qwen2.5-VL accuracy on LVBench by $+10.1$ points. These gains require only $0.11\%$--$0.40\%$ average video-token overhead; analyses further show task-aware allocation and benefits from Adaptive Evidence Invocation.