arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10526cs.AI

智能体不仅会认同,还会记忆:对有状态个人智能体中持续谄媚行为的基准测试

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Self-Improving Personal Agents

Xutao Mao, Liangjie Zhao, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han, Xiang Zheng, Cong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究有状态个人智能体中的持续谄媚行为,引入PASB基准测试,评估真实智能体。通过特定方法隔离写入过程,发现提交边界是关键转折点,提交声明有三种模式,表明智能体谄媚行为是状态写入治理问题,PASB确定相关写入时控制。

中文摘要 AI 辅助

有状态的个人智能体越来越多地维护长期用户档案、情景记忆和可重复使用的技能。这种持续性将对话中的谄媚行为转变为状态写入失败:被接受的以用户为中心的声明可能会被记录为持久偏好、背景事实或工作流程,并在原始对话结束后被再次使用。我们将此称为持续谄媚行为,并引入了个人智能体谄媚行为基准测试(PASB),这是一个包含1600个任务的基准测试,用于追踪对话声明是否被接受、写入持久智能体状态并在后续中立查询中被再次使用。与之前提供预编写记忆的基准测试不同,PASB评估决定存储内容的真实智能体(Hermes-Agent和OpenClaw)。它通过将四种场景框架与四种时间交付模式相结合,并将五轮持久阶段与清除后的三轮查询阶段分开,来隔离写入过程,确保下游影响仅来自持久状态。在十二个模型中,提交边界是关键转折点:下游失败率从仅会话情节中的45.0%增加到提交后的71.9%,一致增加了27.0个百分点。提交的声明呈现出三种写入时模式:状态提升、归因消除和范围扩大。这些模式在类似记忆或程序框架、重复强化甚至跨领域边界的情况下会变得更强。这些结果表明,智能体谄媚行为从根本上说是一个状态写入治理问题。一旦用户内容被提交到持久记忆中,安全性必须控制智能体写入的内容,而不仅仅是它们所说的内容。PASB确定了在保留存储内容的来源、角色和范围的同时,控制有风险提交所需的写入时控制,而不仅仅是在响应级别进行缓解。

英文摘要

Self-improving personal agents now write profiles, memories, and reusable skills that carry over from one chat to the next. Prior work asks whether user pressure bends a model's next answer. Yet these agents can also write the user's claim down, so a later chat may read it back as trusted context. We call this persistent sycophancy. We introduce the Personal Agent Sycophancy Benchmark, PASB, with 1,600 tasks run on two real agents, Hermes-Agent and OpenClaw, across twelve models. Each task isolates a first chat containing the claim from a neutral follow-up chat, so any carryover must pass through a note the agent chose to write. Our analysis shows that downstream failure, meaning how often later answers side with the claim, reason from it, treat it as fact, or stretch it, reaches 71.9% when the follow-up chat can read a saved claim, against 45.0% when the claim stays in the first chat. Writing also edits the claim, as agents save it as a stable preference, a background fact, or a reusable procedure in 51.4% of runs. A saved claim still shapes answers in a different domain. Among the mitigations we test, explicit memory editing helps most, cutting downstream failure to 32.7% for Hermes-Agent and 54.5% for OpenClaw on same-domain follow-ups. PASB shows that self-improving agents must govern what they write down before it governs what they say. Our benchmark is available at https://github.com/henrymao2004/agent-sycophancy.

发表机构

  • City University of Hong Kong(香港城市大学)
  • ByteDance(字节跳动)
  • Yale University(耶鲁大学)
  • Fudan University(复旦大学)
  • Hong Kong Baptist University(香港浸会大学)
  • Dalian University of Technology(大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑