arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

粗粒度索引,细粒度证据:解耦长视频检索增强生成中的时间粒度

Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

Zhe Jin, Zhimin Lin, Bin Zheng, Junhua Fang, Huihua Yang

arXiv 2608.23011首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Soochow University(北京邮电大学; 苏州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长视频RAG中索引与证据粒度耦合的问题,提出无训练的DAGC方法,实现40%-50%节点保留、1.3-1.7倍加速且维持99%问答性能,增益可跨模型迁移。

AI 中文摘要

基于图的检索增强生成(RAG)为长视频理解提供了可扩展的范式,但现有系统在构建检索索引时通常继承视频分割的固定时间粒度。我们认为该设计不必要地将索引粒度与证据粒度耦合:粗粒度表示通常足以定位相关时间区域,而细粒度证据对下游推理仍很重要。我们提出密度感知图构建(Density-Aware Graph Construction,DAGC),一种无训练方法,它将与查询无关的粗粒度检索索引与原始细粒度证据空间解耦。DAGC通过合并视觉上冗余的相邻块构建紧凑、密度自适应的图索引,同时保留与原始时间单元的映射。检索到的粗粒度区域随后被扩展回原始块粒度,用于细粒度证据细化和答案生成。在MLVU、VideoMME和LongVideoBench上的实验表明,DAGC仅保留约40%至50%的原始图节点,实现1.3倍至1.7倍的端到端挂钟加速,同时保留约99%的原始问答性能。该增益可迁移至不同的大型视频语言模型(LVLM)主干和视频RAG管道,表明长视频RAG无需在索引和证据推理中保持相同的时间粒度。

英文摘要

Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing their retrieval index. We argue that this design unnecessarily couples indexing granularity with evidence granularity: coarse representations can often suffice for locating relevant temporal regions, while fine-grained evidence remains important for downstream reasoning. We propose \textbf{Density-Aware Graph Construction (DAGC)}, a training-free approach that decouples a query-independent coarse retrieval index from the original fine-grained evidence space. DAGC constructs a compact, density-adaptive graph index by merging visually redundant neighboring chunks, while preserving mappings to the original temporal units. Retrieved coarse regions are subsequently expanded back to the original chunk granularity for fine-grained evidence refinement and answer generation. Experiments on MLVU, VideoMME, and LongVideoBench show that DAGC retains only about 40--50\% of the original graph nodes and achieves $1.3$--$1.7\times$ end-to-end wall-clock acceleration while preserving approximately 99\% of the original QA performance. The gains transfer across different LVLM backbones and video RAG pipelines, suggesting that long-video RAG need not maintain the same temporal granularity for indexing and evidence reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑