自改进大语言模型智能体中的记忆奖励膨胀
Memory Reward Inflation in Self-Improving LLM Agents
浏览论文内容
中文总结 AI 辅助
该研究发现自改进大语言模型智能体存在记忆奖励膨胀的“回声差距”问题,提出误差独立性假设(EIA),并通过LUCID算法在BIRD文本到SQL基准上提升了执行准确率。
中文摘要 AI 辅助
自改进大语言模型智能体越来越多地从经验中学习,无需更新任何权重。每个回合都被存储在外部记忆中,被评分并检索用于未来类似任务,以塑造后续行为。从奖励角度来看,存储的分数是隐式非参数策略的代理奖励,每个检索到的回合都成为策略改进步骤,其可靠性取决于该分数的产生方式。在部署中,真实标签不可用,因此存储的奖励最多只是大语言模型的评估,这种替换会在研究的基于记忆的自改进智能体和模型系列中产生一种失效模式,即“回声差距”。错误的回合会获得膨胀的奖励,因此智能体会优先重复使用其最确信的错误。由于错误通过记忆而非平均化累积,且确认判断的错误仍与原始自我评分偏差相关,因此无法识别哪些记忆被高估。缺失的属性被形式化为“误差独立性假设(EIA)”,我们证明这是纠正膨胀的必要条件,而非仅对良好验证器的描述:可用信号必须跟踪真实情况,且使其误差与记忆偏差去相关,可恢复的收益恰好是这两个量的闭式函数。我们进一步表明,不仅当检索按存储分数排名时,奖励膨胀会累积,在部署智能体使用的纯相似性检索机制下,奖励膨胀同样会累积。最后,无答案去膨胀算法LUCID在BIRD文本到SQL基准测试中实现了一致的端到端增益,将执行准确率提升至56.9%,高于Memento式自我评分智能体(54.0%,跨随机种子平均增益+2.9个百分点)和相同架构的无记忆智能体(52.4%)。
英文摘要
Self-improving LLM agents increasingly learn from experience without updating any weights. Each episode is stored in an external memory, scored, and retrieved for similar future tasks to shape later behavior. Viewed through a reward lens, the stored score is a proxy reward for an implicit, non-parametric policy. Each retrieved episode then becomes a policy-improvement step whose reliability hinges on how that score is produced. In deployment, ground-truth labels are unavailable, so the stored reward is at best an LLM assessment. This substitution creates a failure mode, the *Echo Gap*, across the memory-based self-improving agents and model families studied. Incorrect episodes receive inflated rewards; thus, the agent preferentially reuses the very mistakes it has most confident in. Because the error compounds through memory rather than averaging out and the confirming judge's errors remain correlated with the original self-grading bias, so it cannot identify which memories are overvalued. The missing property is formalized as the *Error-Independence Assumption* (EIA), which we prove is a *necessary* condition for correcting the inflation, not merely a description of a good verifier: a usable signal must track truth *and* decorrelate its error from the memory bias, and the recoverable payoff is a closed-form function of exactly those two quantities. We further show the inflation compounds not only when retrieval ranks by the stored score but also under plain similarity retrieval which is the regime the deployed agent uses. Finally, the answer-free de-inflation algorithm LUCID delivers a consistent end-to-end gain on the BIRD text-to-SQL benchmark. It raises execution accuracy to $56.9\%$, above both a Memento-style self-graded agent ($54.0\%$, a $+2.9$-point mean gain across seeds) and a memory-less agent of identical architecture ($52.4\%$).
发表机构
- University Of North Texas(北得克萨斯大学)
- Edge Hill University(边山大学)
- University of Cincinnati(辛辛那提大学)
- Iran University of Science and Technology(伊朗科技大学)
- Islamic Azad University(伊斯兰阿扎德大学)
机构由 AI 辅助整理,请以论文原文为准。