发表机构
Meta Reality Labs; University of Virginia; University of Illinois Urbana-Champaign(Meta现实实验室; 弗吉尼亚大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MemLife通过构建实体化第一人称文本情节和时间索引智能体阅读器,实现无需训练的长时程视频记忆推理,在四个基准上提升4.6-12.0%,并借助MemOpt强化学习进一步优化记忆质量。
AI 中文摘要
长时程第一视角视频使个性化AI助手能够对日常生活进行推理。然而,随着视频历史增长到跨越数月或数年的数百小时,为每个查询重新处理原始片段在计算上变得不可行。记忆系统通过将视频压缩为文本表示提供了一种可扩展的替代方案,但在实际基准测试中常常失败:要么记忆未保留关键证据,要么由于搜索空间增长中的检索竞争,检索器无法定位相关条目。为解决这些挑战,我们提出了MemLife,一种多模态记忆系统,它构建基于实体的第一人称文本情节,并通过时间索引的智能体阅读器进行检索。无需训练或查询时视频访问,MemLife在四个长时程基准测试中比最强的无训练基线提高了4.6%至12.0%。为进一步提升记忆质量,我们提出了MemOpt,一种强化学习框架,优化记忆写入器以生成忠实、信息丰富且可检索的记忆。MemOpt在不同视频和问题分布上持续将MemLife提升2.7%至5.0%,其增益可泛化至不同的写入器、阅读器骨干网络及记忆系统。
英文摘要
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.