从错误记忆到修正行动:记忆增强智能体的依赖引导回滚修复
From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents
浏览论文内容
中文总结 AI 辅助
该研究针对记忆增强智能体的错误记忆问题,提出依赖引导回滚修复方法,在150案例受控基准和50案例改编测试中,均优于对比方法,实现了更优的记忆恢复效果与成本权衡。
中文摘要 AI 辅助
持久记忆使基于语言模型的智能体能够跨会话复用信息,但也会让错误持续存在:被污染、过时或属性错误的记录会改变推理、工具使用、回答以及后续的记忆写入。现有防御方法主要是检测或删除可疑记忆,或修正当前响应。删除错误源会使已传播的声明、行动和衍生记忆保持有效,而重置记忆存储或重放完整轨迹会破坏良性状态并重复不必要的计算。因此,我们提出了故障后记忆恢复问题:给定失败的执行过程和已诊断的错误记忆,在保留未受影响工作的同时恢复答案和持久状态。我们的依赖引导回滚修复方法从运行时溯源构建类型化的记忆-行动图,追踪显式下游依赖,保留具有独立可信支持的候选内容,停用无支持的记忆状态,并仅选择性重放与答案相关的受影响计算。我们在涵盖三个工具使用领域和四种记忆错误类型的150个案例的受控基准上,以及在改编自LongMemEval-V2的50个案例轨迹衍生压力测试上评估了该方法。在受控基准上,它实现了85.3%的恢复率,而最佳的竞争恢复方法为77.3%;该方法清除了所有已诊断的错误记忆,保留了所有良性记忆,且仅需选择性重放,LLM调用成本适中。在改编的子集上,它达到了68.0%的恢复率,而次优方法为54.0%,同时还获得了最高的声明失效F1值,为0.669,而次优值为0.603。总体而言,结果并不意味着轨迹重建的全面优化,但表明依赖引导回滚修复在修复错误记忆状态并保留良性记忆的同时,实现了良好的恢复-成本权衡。
英文摘要
Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes. Existing defenses mainly detect or delete suspicious memories, or revise the current response. Deleting the source leaves already propagated claims, actions, and derived memories active, whereas resetting the store or replaying the full trace destroys benign state and repeats unnecessary computation. We therefore formulate \textbf{post-failure memory recovery: } \textit{given a failed execution and diagnosed faulty memories, recover both the answer and persistent state while retaining unaffected work.} Our \textbf{dependency-guided rollback repair} builds a typed memory-to-action graph from runtime provenance, traces explicit downstream dependencies, preserves candidates with independent trusted support, deactivates unsupported memory state, and selectively replays only answer-relevant affected computation. We evaluate this approach on a 150-case controlled benchmark spanning three tool-use domains and four memory failure types, and on a 50-case trajectory-derived stress test adapted from LongMemEval-V2. On the controlled benchmark, it achieves 85.3\% recovery versus 77.3\% for the best competing recovery method, removes all diagnosed faulty memories, preserves all benign memories, and requires only selective replay with modest LLM-call cost. On the adapted subset, it reaches 68.0\% recovery versus 54.0\% for the next best method, while also achieving the highest claim invalidation F1, 0.669 versus 0.603. Overall, the results do not imply uniformly better trace reconstruction, but show that dependency-guided rollback repair provides a strong recovery--cost trade-off while repairing faulty memory state and preserving benign memory.