arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38155cs.CVcs.AIcs.CLcs.IRcs.LG

超越时间线:利用基于实体的传记增强长视频记忆

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于实体的传记框架,通过视觉身份链接分组长视频中同一实体的观察,提升跨事件问答性能,在EgoLifeQA上达72.0%准确率。

中文摘要 AI 辅助

回答关于长视频的问题通常需要跨小时或跨天连接涉及同一对象的事件。按时间顺序的描述和从文本中提取的实体可能无法解决物理身份问题:不同的对象可能共享一个描述,而同一对象的观察结果在不同事件之间仍然不连贯。因此,检索相关事件并不一定能恢复问题所涉及的特定实体的“传记”。为了解决这个问题,我们引入了基于实体的传记(GEB),一种长视频记忆框架,它将跨片段中同一物理实例的视觉基础观察分组为可检索的传记,同时保留每个时刻的上下文。在问答过程中,传记与情节性证据一起被检索,允许模型利用记忆构建期间建立的身份链接,通过事件跟踪实体。在四个基准测试上的评估,包括全天和全周记录,表明在多项选择和开放式问答中,相较于先前的记忆框架有所改进。在EgoLifeQA上,GEB达到了72.0%的准确率,比最佳已发表结果高出4.4个百分点。消融实验表明,基于实体的身份关联和传记阅读都对增益有所贡献,而仅增加描述并不能完全恢复这些增益。

英文摘要

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Amazon.com, Inc.(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

↑