arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExperienceIndex:基于工件的记忆

ExperienceIndex: Artifact-Grounded Memory

Peter Baile Chen, Geoffrey X. Yu, Xinming Liu, Samuel Madden, Dan Roth, Jacob Andreas, Doug Downey, Michael Cafarella

arXiv 2610.10091首次发表:更新:

发表机构

MIT; McKinsey & Company; Oracle AI; UPenn; AI2(麻省理工学院; 麦肯锡公司; 甲骨文人工智能实验室; 宾夕法尼亚大学; 艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ExperienceIndex通过存储单一工件和工件对经验,作为轻量级中间件引导代理找到完整相关工件集,提升答案质量达11.0点并降低在线成本达50.5%,支持跨任务泛化和师生学习。

AI 中文摘要

知识密集型任务需要通过推理共享工件语料库(例如法院案例或科学文献)来回答许多问题。当人类与这些语料库交互时,他们自然地积累关于工件的经验知识,使他们能够快速识别每个新任务的完整相关工件集。然而,现有的AI代理缺乏适当的记忆解决方案来构建或重用这种基于工件的经验,导致答案质量较低和在线成本较高。现有的记忆解决方案从先前的任务解决轨迹中提取和重用信息,但它们主要关注用户偏好、事实属性或抽象推理模式,而不是持久的工件特定知识。我们引入了ExperienceIndex,一种新颖的AI代理经验层,基于先前的推理轨迹捕获和重用关于工件的知识。ExperienceIndex存储两种互补的经验形式:(i)单一工件经验,总结工件对先前任务的贡献;(ii)工件对经验,编码在过去推理中发现的结构关系。作为轻量级中间件集成,ExperienceIndex使用经验检索机制引导代理找到新任务的完整相关工件集,提高答案质量和效率。在不同的语料库和具有不同搜索框架的代理解决方案中,ExperienceIndex带来一致的改进,将答案质量提高多达11.0个百分点,并将在线美元成本降低多达50.5%。我们进一步展示了两个好处:(i)跨任务泛化,从文本到SQL任务积累的经验转移到同一工件语料库上的事实问答任务;(ii)师生学习,来自更强模型的经验使较弱模型达到相当的性能。

英文摘要

Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solving traces, but they primarily focus on user preferences, factual attributes, or abstract reasoning patterns rather than persistent artifact-specific knowledge. We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces. ExperienceIndex stores two complementary forms of experience: (i) single-artifact experiences that summarize an artifact's contribution to prior tasks and (ii) artifact-pair experiences that encode structural relationships discovered during past reasoning. Integrated as lightweight middleware, ExperienceIndex uses an experience retrieval mechanism to guide agents toward the complete set of relevant artifacts for new tasks, improving both answer quality and efficiency. Across diverse corpora and agentic solutions with different search frameworks, ExperienceIndex delivers consistent gains, raising answer quality by up to 11.0 points and reducing online dollar cost by up to 50.5%. We further demonstrate two benefits: (i) cross-task generalization, where experiences accumulated from text-to-SQL tasks transfer to factoid QA tasks over the same artifact corpus, and (ii) teacher-student learning, where experiences from a stronger model enable a weaker model to reach comparable performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑