发表机构
Johns Hopkins University; Cornell University(约翰·霍普金斯大学; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能代理处理冲突记忆的问题,提出E-P-R框架,通过在WebArena和MemTrapBench上实例化发现主要失败始于入口,冲突记忆导致合规陷阱,强调评估记忆增强型代理应考虑其处理记忆的全过程。
AI 中文摘要
记忆正成为长期人工智能代理的核心组件,现有工作多将记忆视为供应问题。但我们仍缺乏对模型在多步行动轨迹中如何使用检索到的记忆的清晰认识。为此提出轨迹级框架E-P-R来诊断此过程,在WebArena和MemTrapBench上进行实例化。发现主要失败常始于入口,冲突记忆会导致合规陷阱,各模型的合规率相似,但合规后成功率极低。结果表明,对记忆增强型代理的评估不仅应基于检索质量或最终成功率,还应考虑其在整个轨迹中如何处理记忆。
英文摘要
Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve for the next task. Yet we still lack a clear account of how models consume retrieved memory across a multi-step action trajectory. This consumption process matters because it determines not only what memories should be retrieved, but also what models and control policies are needed to use them safely. To diagnose this process, we propose Entry--Propagation--Recovery (E-P-R), a trajectory-level framework that asks where memory first changes an action, whether that change carries forward, and whether the agent can recover after leaving a correct path. We instantiate E-P-R on WebArena and on MemTrapBench, a controlled benchmark we build to isolate these phases. We find that the main failure often begins at entry: agents adopt conflicting memory at the first exposed decision point even when it is task-wrong. Repeated exposure then amplifies this early error, while recovery after divergence is weak. Together, these effects create a compliance trap: across models, conflicting memory induces similar compliance rates, but once agents comply, their success rates collapse to a low floor. Stronger agents therefore suffer larger absolute damage because each compliance event erases more baseline capability. These results suggest that memory-augmented agents should be evaluated not only by retrieval quality or final success rate, but by how they consume memory throughout the trajectory.