arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12627cs.CVcs.AIcs.CLcs.HC

EgoCITE:面向长时程自我中心记忆的上下文增强索引与时序感知检索

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

Le Zhang, Hao Chen, Vlad Roznyatovskiy, Jianzhong Zhang, Ke Sun

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对长时程自我中心记忆系统的索引不可靠、忽略时序意图的问题,提出EgoCITE框架,经多数据集评估,其准确率优于基线且成本显著低于长上下文LLM智能体。

中文摘要 AI 辅助

长时程自我中心记忆将连续的第一人称视频和音频转化为可检索的过往经验记录。我们发现现有系统存在两个瓶颈:基于缺乏上下文的字幕构建的索引对于智能体搜索不可靠,同时检索过程忽略了问题的时序意图。为解决这两个瓶颈,我们提出EgoCITE(Egocentric Context-augmented Indexing and Time-aware Evidence Retrieval,自我中心上下文增强索引与时序感知证据检索),这是一个用于自我中心问答(QA)的长时程智能体记忆框架。EgoCITE包含三个组件:EgoScheme利用局部多模态上下文将碎片化的视频字幕和语音转录转化为独立的原子记忆索引;EgoIndex将互补的动作、活动、话语和对话表示组织为多粒度的可检索多视图记忆索引;EgoRetrv将语义搜索与问题条件化的时序相关性评分及检索证据筛选相结合。我们在EgoLifeQA、EgoMem和EgoR1-Bench数据集上,以答案准确率和目标事件检索一致性为指标评估EgoCITE,结果显示EgoCITE相较于智能体记忆基线的准确率提升了至少4.4%至14.2%,同时成本比长上下文大语言模型(LLM)智能体低36倍。

英文摘要

Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCITE comprises three components. EgoScheme uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices. EgoIndex organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities. EgoRetrv combines semantic search with question-conditioned temporal relevance scoring and curation of retrieved evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment. EgoCITE improves accuracy over agentic memory baselines by at least 4.4--14.2% while achieving 36$\times$ lower cost than long-context LLM agents.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

↑