发表机构
MIIT Key Laboratory of Data and Decision Intelligence; Beihang University(工业和信息化部数据与决策智能重点实验室; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体记忆检索中反馈成本高的问题,提出UpliftMem,通过集合级执行提升学习检索,利用EVSI准则分配探测,在多个基准上取得最优成功率。
AI 中文摘要
大型语言模型(LLM)智能体复用外部记忆来指导新任务,但有效的检索需要学习哪些记忆集合能改善执行。这种学习依赖于代价高昂的结果反馈:普通检索仅观察已执行的集合,而评估替代方案则需要额外的模拟运行。我们引入了\ extsc{UpliftMem},它从相对于无记忆的同一执行器的集合级执行提升中学习记忆检索。关于检索偏好如何限制反馈覆盖范围的理论分析,促使对替代记忆集合进行有针对性的探测。探测选择遵循样本信息期望值(EVSI)准则,该准则在相关高斯模型下以闭式推导得出,根据有限的训练模拟运行对局部检索决策的预期改进来分配这些运行。共享评分器使用冻结的执行器进行训练,并在测试时无需探测即可选择记忆集合。在ALFWorld、WebShop和BigCodeBench上,\ extsc{UpliftMem}在主要评估集上取得了评估基线中最佳的成功率。受控的固定存储和匹配探测预算评估进一步证明了改进的记忆使用决策和更有效地利用执行反馈。
英文摘要
Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textsc{UpliftMem} achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.