arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoMonth:面向长期时空记忆的月级自我中心视频基准

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi

arXiv 2608.13113首次发表:更新:

发表机构

Nanjing University; Huawei Technologies Co., Ltd.; Tianjin University(南京大学; 华为技术有限公司; 天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员推出首个月级自我中心视频基准EgoMonth,评估发现当前主流MLLMs的长期时空记忆能力远逊于人类,仅为有损总结器而非忠实记忆器。

AI 中文摘要

多模态大语言模型(MLLMs)的最新进展推动了视频理解领域的长足进步,随之涌现出越来越多的长视频基准。然而,现有基准大多依赖网络来源的视频,这类视频缺乏片段间的时空连续性,难以评估模型能否在数天或数周的真实世界体验中保持一致的记忆。我们推出EgoMonth,这是首个月级自我中心视频理解基准。EgoMonth包含来自20名参与者、时长跨度20至120天的300多小时第一人称日常生活记录,搭配1443个人工构建的多项选择题-答案对。我们设计了一个基于认知原理的14任务评估框架,该框架分为三个层级的认知水平:图式巩固、情景索引和级联推理。对最先进的开源及闭源MLLMs的评估显示,即使是表现最佳的模型Gemini 2.5 Pro,也仅达到71.8%的宏平均准确率,比修正后的人类基准94.2%低22.4个百分点。部分模型在路径推理、跨视图空间推理和方向判断等任务上的表现接近或低于25%的随机水平,而即使是最强的闭源模型,其表现也远低于人类水平。这些结果表明,当前的MLLMs更像是有损总结器,而非忠实记忆器,凸显了对具备真正长期时空记忆能力的架构的需求。

英文摘要

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.

Comments21 pages, 4 figures, 6 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑