发表机构
Nanjing University; Huawei Technologies Co., Ltd.; Tianjin University(南京大学; 华为技术有限公司; 天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员推出首个月级自我中心视频基准EgoMonth,评估发现当前主流MLLMs的长期时空记忆能力远逊于人类,仅为有损总结器而非忠实记忆器。
AI 中文摘要
多模态大语言模型(MLLMs)的最新进展推动了视频理解领域的长足进步,随之涌现出越来越多的长视频基准。然而,现有基准大多依赖网络来源的视频,这类视频缺乏片段间的时空连续性,难以评估模型能否在数天或数周的真实世界体验中保持一致的记忆。我们推出EgoMonth,这是首个月级自我中心视频理解基准。EgoMonth包含来自20名参与者、时长跨度20至120天的300多小时第一人称日常生活记录,搭配1443个人工构建的多项选择题-答案对。我们设计了一个基于认知原理的14任务评估框架,该框架分为三个层级的认知水平:图式巩固、情景索引和级联推理。对最先进的开源及闭源MLLMs的评估显示,即使是表现最佳的模型Gemini 2.5 Pro,也仅达到71.8%的宏平均准确率,比修正后的人类基准94.2%低22.4个百分点。部分模型在路径推理、跨视图空间推理和方向判断等任务上的表现接近或低于25%的随机水平,而即使是最强的闭源模型,其表现也远低于人类水平。这些结果表明,当前的MLLMs更像是有损总结器,而非忠实记忆器,凸显了对具备真正长期时空记忆能力的架构的需求。
英文摘要
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.
Comments21 pages, 4 figures, 6 tables, including appendices