arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越视觉边界:为电影检索增强生成(RAG)重新思考场景分割

Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG

Dong-Hee Kim, Seonwoo Choi, Changbeen Kim, Jungmyung Wi, Juyeon Ko, Youngju Choi, Il Hyeon Mun, Hyunwoo J. Kim, Donghyun Kim

arXiv 2608.28699首次发表:更新:

发表机构

Korea University; KAIST(高丽大学; 韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对电影RAG场景,发现现有场景分割方法不如均匀时间分块,提出以叙事为核心的NarraScene数据集,其生成的片段作为检索单元可提升下游电影理解任务性能。

AI 中文摘要

理解长视频内容仍是多模态大语言模型(MLLM)面临的核心挑战:稀疏帧采样无法捕捉细粒度视觉细节,而密集采样则会快速超出上下文长度限制。检索增强生成(RAG)提供了一种有前景的折中方案,通过选择性检索相关视频片段来支撑生成,但其效果关键取决于作为检索单元的视频片段质量。本文针对面向电影理解的RAG展开研究,该任务需要对跨越数小时内容的角色、事件及叙事弧进行故事级推理。场景分割作为一项长期研究的问题,可将电影划分为语义连贯的单元,是定义此类检索单元的自然候选方案。我们通过对下游电影理解任务的综合评估,重新审视现有方法是否真正适用于该场景,发现它们始终未能优于简单的均匀时间分块。我们对最标准的场景分割基准的分析揭示了原因:当前标注优先考虑视觉显著的转换,而非叙事事件结构。针对这种不匹配,我们引入了NarraScene,这是一个以叙事为核心的场景分割数据集,采用涵盖物理、角色和叙事变化的三级认知分类法进行标注,其中每个有效边界都需要叙事层面的转变。当将这些基于叙事的片段用作检索单元时,它们在下游电影理解任务上的表现优于均匀分块,这表明电影RAG中场景分割的核心挑战并非检测边界,而是识别对电影理解重要的叙事事件单元。

英文摘要

Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, while dense sampling quickly exceeds context length limits. Retrieval-augmented generation (RAG) offers a promising middle ground by selectively retrieving relevant video segments for grounded generation, yet its effectiveness critically depends on the quality of the video segments used as retrieval units. In this paper, we investigate RAG for movie understanding, which demands story-level reasoning over characters, events, and narrative arcs spanning hours of content. Scene segmentation, a long-studied problem that partitions movies into semantically coherent units, is a natural candidate for defining such retrieval units. We reexamine whether existing methods actually serve this role through comprehensive evaluation on downstream movie understanding tasks, and find that they consistently fail to outperform naive uniform temporal chunking. Our audit of the most standard scene segmentation benchmarks reveals why: current annotations prioritize visually salient transitions over narrative event structure. Motivated by this mismatch, we introduce NarraScene, a narrative-centric scene segmentation dataset annotated with a three-level cognitive taxonomy spanning physical, character, and narrative change, where every valid boundary requires a narrative-level shift. When used as retrieval units, these narrative-grounded segments outperform uniform chunking on downstream movie understanding tasks, suggesting that the central challenge for scene segmentation in movie RAG is not detecting boundaries, but identifying the narrative event units that matter for movie understanding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑