arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIDGE:用于长视频理解的区域感知导数引导证据选择

RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding

Shanqing Xu, Meng Luo, Mengchen Qian, Yuhui Gao, Siyue Peng, Xiaohan Zhong, Xiaojin Zhang, Zhongyu Wei, Wei Chen, Xiang Bai

arXiv 2608.29958首次发表:更新:

发表机构

Huazhong University of Science and Technology; National University of Singapore; Fudan University(华中科技大学; 新加坡国立大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频理解中视觉内容超出LVLM固定token预算的问题,提出RIDGE框架,通过将帧-查询相似性曲线作为时间信号划分区域并选择证据,在多基准和主干上取得最优性能。

AI 中文摘要

长视频包含的视觉内容远多于大型视觉语言模型(LVLMs)在固定视觉token预算下可处理的内容,因此帧选择至关重要。现有的查询感知选择器通常会估计帧与查询的相关性,并从高分帧中构建紧凑子集。尽管它们的机制不同,但相似性序列仍常被主要视为用于排名或采样的值,而非反映与查询相关的证据如何随时间出现、达到峰值和消失的有序信号。这可能会掩盖那些解释事件、为事件提供背景或紧随事件之后的帧,因为此类证据可能位于附近相关性峰值的上升或下降侧,从而获得较低的绝对分数。我们提出RIDGE,这是一个将帧与查询的相似性曲线作为时间信号读取的帧选择框架。通过利用局部变化和曲率,RIDGE将时间线划分为结构区域,并应用特定于区域的选择,以在固定预算下保留事件核心、过渡、积累、后续影响和上下文帧。它是对预先计算的帧-查询分数的轻量级后处理步骤,既不需要训练,也不需要迭代调用LVLMs。在四个长视频基准和三个主干网络上,RIDGE在大多数设置中取得了最佳性能,在其他设置中也保持了竞争力。

英文摘要

Long videos contain far more visual content than Large Vision-Language Models (LVLMs) can process under a fixed visual-token budget, making frame selection essential. Existing query-aware selectors usually estimate frame-query relevance and build a compact subset from high-scoring frames. Although their mechanisms differ, the similarity sequence is still often treated primarily as values to rank or sample from, rather than as an ordered signal whose shape reflects how query-relevant evidence emerges, peaks, and fades over time. This can obscure frames that explain, contextualize, or follow an event, because such evidence may lie on the rising or falling sides of a nearby relevance peak and receive lower absolute scores. We propose RIDGE, a frame selection framework that reads the frame-query similarity curve as a temporal signal. By using local changes and curvature, RIDGE partitions the timeline into structural regions and applies region-specific selection to preserve event cores, transitions, buildup, aftermath, and contextual frames under a fixed budget. It is a lightweight post-processing step on precomputed frame-query scores and requires neither training nor iterative LVLM calls. Across four long-video benchmarks and three backbones, RIDGE achieves the best performance in most settings and remains competitive in the others.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑