arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35596cs.CRcs.AIcs.CL

SEABench:自我进化智能体中的内生错位基准测试

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi, Ferdinando Fioretto

首次发表
浏览论文内容

中文总结 AI 辅助

SEABench通过48个纵向任务序列基准测试自我进化LLM智能体的内生错位,发现进化提升任务完成率但引发安全失败,并利用思维链推理实现低误报率监控。

中文摘要 AI 辅助

自我进化的LLM智能体因其在部署后通过修改其控制框架(包括控制器指令、内存管理协议以及可复用的工具和技能)来响应用户和环境反馈、从而提升自身能力的能力而受到广泛关注。然而,局部有用的更新可能会持续到后续任务中,在这些任务中产生不安全的行为,即使没有直接的对抗性影响。为了研究这一风险,我们引入了SEABench,一个用于研究智能体自我进化所引发的内生错位的基准测试,包含48个纵向任务序列,覆盖了丰富的个人助理环境中的多个进化表面、任务领域和危害类型。为了考虑智能体操作中固有的随机性,我们提供了一个自适应轨迹发现流程,该流程在保持原始任务意图的同时探测失败,并通过配对的非进化智能体和归因分数支持因果归因。我们在多个最近的LLM、进化表面和危害类型上的评估表明,自我进化确实提高了任务完成率,但往往以安全失败为代价,而这些失败在配对的非进化基线智能体中是不存在的。我们还表明,在不同的进化表面和危害类型中,会出现定性不同的安全行为。此外,我们表明,这种安全行为的差异反映在智能体的思维链推理中,这提供了一种有效的监控策略,可以以低误报率缓解不安全行为。

英文摘要

Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operations, we provide an adaptive trajectory discovery pipeline that probes for failures while preserving original task intent and supports causal attribution through paired non-evolving agents and attribution scores. Our evaluation across multiple recent LLMs, evolution surfaces, and harm types reveals that self-evolution indeed increases task completion rates but often at the cost of safety failures that are absent for paired non-evolving baseline agents. We also show that qualitatively different safety behaviors emerge across evolution surfaces and harm types. Further, we show that this divergence in safety behavior is reflected in agents' chain-of-thought reasoning, which yields an effective monitoring strategy that can mitigate unsafe behavior with a low false positive rate.

发表机构

  • University of Virginia(弗吉尼亚大学)
  • ELLIS Institute Tübingen(图宾根ELLIS研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑