发表机构
Nankai University; National University of Singapore; Shanghai Institute of Optics and Fine Mechanics, Chinese Academy of Sciences; Peking University; Technical University of Munich; Queen Mary University of London; Wuhan University; University of New South Wales(南开大学; 新加坡国立大学; 中国科学院上海光学精密机械研究所; 北京大学; 慕尼黑工业大学; 伦敦大学玛丽女王学院; 武汉大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体记忆的不可靠准入与漂移问题,提出MemGuard方法,通过保留验证器的生命周期元数据提升多任务基准的成功指标与效率,效果优于现有基线。
AI 中文摘要
大语言模型(LLM)智能体正从单一提示词使用转向长任务流应用,其中可复用记忆成为终端、软件工程及网页任务的核心能力。这类记忆仅在存储的经验能在数百次交互中保持可靠时才有用,但实际存在两种故障模式打破该假设:第一种是不可靠准入,失败轨迹、偶然成功及误导性观察因看似相关进入记忆,进而误导后续决策;第二种是记忆漂移,长期运行的记忆库会积累重复、过时及冲突记录,仅靠检索无法修复。MemGuard的核心区别在于将验证器输出视为持续的生命周期元数据而非一次性过滤器,它将多标准分数-标记验证转换为奖励、置信度、标签及不确定性描述符,这些描述符在激活前附加到每个候选项,并在检索、冲突解决、摘要生成及归档过程中复用。我们在四个主干模型上针对Terminal-Bench 2.0、SWE-Bench Verified、WebArena及Mind2Web四个基准评估MemGuard,在匹配的运行时预算下与四个记忆基线及仅验证器对照组对比。在5个随机种子的平均结果中,MemGuard在全部16个主干-基准设置中取得最佳成功指标和最低平均步骤数,较我们评估的记忆方法中最强的现有基线ReasoningBank有所提升,在WebArena上的最大提升为7.9个成功率百分点,在Mind2Web上为5.6个步骤成功率百分点,在终端及软件工程基准上为2.4-3.5个百分点。代码可在该https URL获取。
英文摘要
LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering, and web tasks. Such memory is useful only when stored experience remains reliable across hundreds of interactions, but two failure modes break that assumption in practice. The first is unreliable admission: failed trajectories,accidental successes, and misleading observations enter memory because they appear relevant, then mislead later decisions. The second is memory drift: long-running banks accumulate duplicate, stale, and conflicting records that retrieval alone cannot repair. MemGuard's key distinction is to treat verifier output not as a one-shot filter, but as persistent lifecycle metadata. It converts multi-criteria score-token verification into reward, confidence, label, and uncertainty descriptors that are attached to every candidate before activation and reused during retrieval, conflict resolution, summarization, and archival. We evaluate MemGuard on Terminal-Bench 2.0, SWE-Bench Verified, WebArena, and Mind2Web across four backbones, comparing against four memory baselines plus a verifier-only control under matched runtime budgets. Averaged over five seeds, MemGuard achieves the best success metric and lowest average steps in all 16 backbone-benchmark settings, improving over ReasoningBank, the strongest prior baseline among the memory methods we evaluate, with a largest gain of 7.9 success-rate points on WebArena, 5.6 step-success-rate points on Mind2Web, and 2.4-3.5 points on terminal and software-engineering benchmarks. Code is available at https://github.com/whyyyyy123/MemGuard.
Comments30 pages, 7 figures