发表机构
School of Cyber Science and Technology, Shandong University; State Key Laboratory of Cryptography and Digital Economy Security, Shandong University; Shandong Key Laboratory of Artificial Intelligence Security, Shandong University(山东大学网络空间安全学院; 山东大学密码学与数字经济安全国家重点实验室; 山东大学人工智能安全山东省重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可控框架\textsc{MemGauge},在11个LLM和两个记忆基准上揭示LLM智能体记忆各阶段的效用-风险权衡特征,为阶段感知评估与控制提供依据。
AI 中文摘要
长期记忆正成为大型语言模型(LLM)智能体的核心能力,可实现个性化和长周期交互。然而,保留、转换或暴露更多信息的记忆机制,既会影响良性效用,也会增加记忆投毒的易感性。现有评估通常在固定配置下单独测量记忆效用或攻击风险,对阶段特定设计选择如何重塑二者权衡的洞见有限。本文提出\textsc{MemGauge},这是一个可控框架,可在匹配的干净和投毒条件下分别调整写入准入、管理策略和检索暴露。在11个LLM和两个长期记忆基准上的受控评估揭示了三种不同特征:写入期间类似阈值的风险转变、管理期间依赖策略的局部解耦,以及检索期间效用与风险的耦合增长。我们进一步将类似的阶段级测量应用于四个现有记忆系统,观察到与这些特征定性一致的诊断关联。这些结果表明,针对性投毒风险因记忆操作而异,并推动对LLM智能体记忆的阶段感知评估与控制。
英文摘要
Long-term memory is becoming a core capability of LLM agents, enabling personalization and long-horizon interaction. However, memory mechanisms that retain, transform, or expose more information can affect both benign utility and susceptibility to memory poisoning. Existing evaluations typically measure memory utility or attack risk in isolation under fixed configurations, providing limited insight into how stage-specific design choices reshape their trade-off. We present \textsc{MemGauge}, a controllable framework that separately varies writing admission, management policy, and retrieval exposure under matched clean and poisoned conditions. Across 11 LLMs and two long-term memory benchmarks, controlled evaluations reveal three distinct profiles: a threshold-like risk transition during writing, policy-dependent local decoupling during management, and coupled growth of utility and risk during retrieval. We further apply analogous stage-level measurements to four existing memory systems and observe diagnostic associations qualitatively consistent with these profiles. These results show that targeted poisoning risk varies across memory operations and motivate stage-aware evaluation and control of LLM-agent memory.