发表机构
University of Science and Technology of China; Jilin University; Hangzhou Dianzi University; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(中国科学技术大学; 吉林大学; 杭州电子科技大学; 合肥综合性国家科学中心人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VideoTapestry提出无需训练的多智能体框架,通过查询驱动的从粗到细精炼预构建分层视频记忆,在多个长视频基准上显著超越GPT-5.5,实现最先进性能。
AI 中文摘要
长视频理解对记忆提出了很高的要求,因为回答问题往往需要检索分布在长时间跨度上的信息。现有方法大致遵循两种范式:查询驱动的探索,该方法对定位误差敏感;以及查询无关的记忆构建,该方法可能遗漏特定于问题的细节。我们提出了VideoTapestry,一个无需训练的多智能体框架,通过从粗到细、查询驱动的精炼过程,调整预先构建的分层视频记忆。预先构建的记忆将视频内容组织为三个层次,分别捕捉全局叙事上下文、事件级时间结构和细粒度关系证据。为了支持从粗到细的定位和观察,我们为每个层次分配一个专门的智能体,将检索和精炼保持在特定尺度的上下文中。在查询的引导下,这些智能体重新访问相关的视频区域,并用有针对性的多模态观察来丰富各层记忆。它们的精炼结果根据原始层次结构组装成一个复合的查询自适应记忆,以紧凑的形式保留全局上下文,同时沿查询相关分支保留细粒度证据以供最终推理。与直接的GPT-5.5推理相比,VideoTapestry在LVBench、LongVideoBench(Long)、Video-MME(Long)和EgoSchema上分别实现了17.2%、14.9%、9.8%和7.0%的绝对准确率提升,在所有竞争者中取得了最先进的结果。
英文摘要
Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensitive to localization errors, and query-independent memory construction, which may omit question-specific details. We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. The preconstructed memory organizes video content into three levels, capturing global narrative context, event-level temporal structure, and fine-grained relational evidence, respectively. To support coarse-to-fine localization and observation, we assign a specialized agent to each level, keeping retrieval and refinement within a scale-specific context. Guided by the query, these agents revisit relevant video regions and enrich layer-wise memories with targeted multimodal observations. Their refinements are assembled according to the original hierarchy into a composite query-adaptive memory, preserving global context in a compact form while retaining fine-grained evidence along query-relevant branches for final reasoning. Compared with direct GPT-5.5 inference, VideoTapestry achieves absolute accuracy gains of 17.2%, 14.9%, 9.8%, and 7.0% on LVBench, LongVideoBench (Long), Video-MME (Long), and EgoSchema, respectively, achieving the state-of-the-art results among all competitors.