arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31874cs.AI

COUNTERMEM:语言智能体的世界模型验证反事实记忆

COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents

Hongji Pu, Ruixiang Tang, Yongfeng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

COUNTERMEM提出一种强化学习框架,利用世界模型验证反事实记忆,提升语言智能体跨任务性能,在12个基准上平均提升12.6个百分点。

中文摘要 AI 辅助

现有的智能体记忆框架主要通过智能体与事实世界的交互来创建记忆,例如记住所采取行动的反馈以提升未来任务的表现。然而,这些框架在记忆构建过程中很少提出“如果”的问题:如果采取了不同的行动,反馈是否会改变,以及这种反馈如何成为有用的记忆?在活跃环境中直接获取此类反馈可能代价高昂,并且可能改变用于比较所需的状态。在这项工作中,我们引入了COUNTERMEM,一个用于跨任务构建和使用经过验证的反事实记忆的强化学习框架。在一次失败的行动之后,COUNTERMEM使用可执行的世界模型(如测试、证明检查器和求解器)从原始状态的副本或重置中评估局部替代方案。它存储改进措施以及原始和纠正后的行动、检查结果和重用条件。一个学习到的记忆使用策略选择检索到的记录或跳过记忆,以平衡任务成功率和交互成本,而基础LLM保持不变。在保留评估期间,记忆和策略都被冻结。我们在六个领域的12个基准设置上评估了COUNTERMEM。使用gpt-oss-120b,COUNTERMEM在所有12个跨六个领域的基准上同时改进了ReAct和Reflexion,相对于其未增强版本平均提高了12.6个百分点。在两个骨干网络的四领域比较中,任务运行令牌减少了7.7%至42.0%,不包括离线选择器训练成本。进一步的分析表明,移除验证或持久存储会削弱收益,而将经过验证的纠正应用于不合适的决策可能会逆转这些收益。代码将在论文被接受后发布。

英文摘要

Existing agent memory frameworks mainly create memory through an agent's interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the "what if" question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.

补充信息

↑