记住、验证还是询问?大语言模型智能体中记忆承诺的跨家族评估
Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究通过MCB数据集评估LLM智能体的记忆承诺,发现模型验证事实变化更可靠,策略提示可降低错误持久率,少样本提示能提升部分任务准确率,但澄清召回率仍较低。
中文摘要 AI 辅助
持久记忆可使大语言模型(LLM)智能体实现个性化,但错误的持久更新会悄无声息地扭曲未来行为。本研究探讨记忆澄清边界:源自交互的信息应被持久存储、仅用于当前上下文、重新验证,还是向用户澄清。MCB数据集包含140个主要场景,分为70个开发项和70个保留项,另有一个独立的70项对比集,该数据集同时评估动作标签和结构化工具调用选择。两名非作者独立对70个保留主要项和70个对比项进行标注,标注一致性达97.1%,Cohen's kappa值为0.962;一名盲法第三方解决4处分歧,以非作者多数意见替换8处作者标注。在Claude和Qwen模型上的实验显示,模型验证变化事实的可靠性高于请求用户解决歧义的可靠性。仅使用Qwen模型时,在12个澄清项中弃权(不执行)0次,在18个新鲜度项中验证12次。少样本提示将准确率从0.557提升至0.771(配对差值为+0.214,经Holm调整的精确McNemar检验p值p_H=0.002),但澄清召回率仍为0.333。策略提示将错误持久率从0.243降至0.100(p_H=0.038),不过其准确率提升并不显著。每个Claude模型的标注-工具一致性为57%,Qwen模型仅为23%;Qwen准确率从0.557降至0.343(p_H=0.047)。研究指出,记忆评估必须同时测试明确决策和工具调用选择。
英文摘要
Persistent memory can personalize an LLM agent, but an incorrect durable update can silently distort future behavior. We study the memory-clarification boundary: whether interaction-derived information should be persisted, used only in the current context, re-verified, or clarified with the user. MCB contains 140 primary scenarios, split into 70 development and 70 held-out items, plus a separate 70-item contrast set. It evaluates both action labels and structured tool-call selection. Two non-authors independently label the 70 held-out primary and 70 contrast items (97.1% agreement, Cohen's kappa = 0.962); a blind third resolves four disagreements, replacing eight author labels by non-author majority. Across Claude and Qwen, models verify changing facts more reliably than they ask users to resolve ambiguity. Bare Qwen asks on 0/12 clarification items while verifying 12/18 freshness items. Few-shot prompting raises accuracy from 0.557 to 0.771 (paired delta = +0.214, Holm-adjusted exact McNemar p_H = 0.002), yet clarification recall remains 0.333. The policy prompt reduces erroneous persistence from 0.243 to 0.100 (p_H = 0.038), although its accuracy gain is not significant. Label-tool agreement is 57% for each Claude model and 23% for Qwen; Qwen accuracy falls from 0.557 to 0.343 (p_H = 0.047). Memory evaluation must test both stated decisions and tool-call choices.
发表机构
- Southern Methodist University(南卫理公会大学)
- Washington University in St. Louis(圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。