arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

记忆何时有用?工具使用型LLM智能体中长期记忆的成本感知评估

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

Shweta Mishra, Shashank Mishra

arXiv 2609.05441首次发表:更新:

发表机构

Independent Research(独立研究)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MERIT基准,在成本核算下评估工具使用型LLM智能体的长期记忆边际效用,发现记忆显著提升任务成功率,但更新事实下嵌入检索不稳定,全回放不经济。

AI 中文摘要

当前对LLM智能体长期记忆的评估依赖于对话回忆基准(LoCoMo、LongMemEval),这些基准衡量的是基于对话历史的问题回答能力,而非记忆的事实是否改变了工具使用型智能体的行为。我们提出了MERIT(面向现实仪器化任务的记忆评估),这是一个基准测试与测试框架,用于在显式成本核算下衡量记忆对任务执行型智能体的边际效用。MERIT提供了三个领域的 episodic 工具使用任务,其任务对早期片段事实的依赖通过自动化泄漏检查得到验证;难度阶梯以更新事实回忆结束;受控记忆损坏;以及对每次记忆操作的完整 token 和美元计量。在23,440个计分片段($42.57)中,基于gpt-4.1-mini的两代试点和预注册的3模型×3种子网格(GPT-4.1、Claude Haiku 4.5;记忆端固定),记忆将依赖任务的成功率从泄漏验证的基线0.00提升至0.55-1.00。在更新事实方面,嵌入检索的表现不可预测地崩溃(各模型0.30-0.95;最大种子差距0.45),且智能体仅在55%的情况下基于正确检索到的值采取行动,而写入时更新存储(结构化事实存储,尤其是LLM摘要)保持在0.70-1.00;混合方案比单独的事实存储更差。最新一代抽查(Claude Sonnet 5,以干净的全回放控制为门控)重现了这一模式。更换记忆的实现会使任务成功率变化高达60个百分点,且全回放从不经济:每个领域的最佳条件每美元提供2.7-3.9倍的边际效用。我们发布了该基准、测试框架及所有轨迹。

英文摘要

Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting. MERIT provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check; a difficulty ladder ending in updated-fact recall; controlled memory corruption; and full token and dollar metering of every memory operation. Across 23,440 scored episodes ($42.57), a two-generation pilot on gpt-4.1-mini and a preregistered 3-model x 3-seed grid (GPT-4.1, Claude Haiku 4.5; memory side held fixed), memory lifts dependent-task success from a leak-verified floor of 0.00 to 0.55-1.00. On updated facts, embedding retrieval collapses unpredictably (0.30-0.95 across models; max seed gap 0.45), and agents act on a correctly retrieved value only 55% of the time, while update-on-write stores (a structured fact store and, notably, LLM summarization) remain at 0.70-1.00; the hybrid is worse than the fact store alone. A latest-generation spot-check (Claude Sonnet 5, gated on a clean full-replay control) reproduces the pattern. Swapping a memory's implementation moves task success by up to 60 points, and full replay is never economical: the best condition per domain delivers 2.7-3.9x its marginal utility per dollar. We release the benchmark, harness, and all traces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑