arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自进化智能体的安全性:综述

Safety in Self-Evolving Agents: A Survey

Jiahao Chen, Zhou Feng, Oubo Ma, Yichen Yan, Ruixiao Lin, Hangtao Zhang, Linkang Du, Hengyu An, Yong Yang, Jun Liu, Junhao Li, Naen Xu, Chunyi Zhou, Yuan Su, Zehao Jin, Qianli Ma, Leyi Qi, Yiming Wang, Zhe Ma, Yuwen Pu, Mengyao Du, Yuanyi Song, Enhao Huang, Zhihui Fu, Jun Wang, Jinfeng Li, Yuefeng Chen, Hui Xue, Yiming Li, Tianyu Du, Shouling Ji

arXiv 2610.00093首次发表:更新:

发表机构

Zhejiang University; Huazhong University of Science and Technology; Xi’an Jiaotong University; National University of Defense Technology; Shanghai Jiao Tong University; Georgia Institute of Technology; University of Science and Technology of China; Nanyang Technological University; OPPO Research Institute; Tianjin University; Chongqing University; Alibaba Group; Rakuten Group(浙江大学; 华中科技大学; 西安交通大学; 国防科技大学; 上海交通大学; 佐治亚理工学院; 中国科学技术大学; 南洋理工大学; OPPO研究院; 天津大学; 重庆大学; 阿里巴巴集团; 乐天集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述提出SAVER框架,系统分析自进化智能体因经验复用导致的安全风险,强调纵向评估以追踪不安全影响的起源、修复及再涌现。

AI 中文摘要

大型语言模型(LLMs)展现出强大的通用能力,但其参数在部署后通常保持固定,限制了从新交互中学习的能力。在开放环境中,这促使了自进化智能体的出现,它们持续从数据、反馈和积累的经验中更新可复用的状态——包括模型参数、记忆、工具定义、技能和工作流。这一转变改变了安全问题:一旦经验成为可复用的状态,过去的事件就变成未来的原因,在一种情境下无害的信息可能在后来的决策中产生更大的持久性、权威性或范围的影响。因此,自进化智能体的安全性不仅询问一个响应是否对齐或一个动作是否被授权,还询问安全属性是否能在局部有用经验的积累、泛化和跨情境复用中得以保持。我们引入了SAVER,一个以转变为中心的框架,其中Substrate定位可复用的影响,Adaptation捕捉其变化方式,Violation识别受损的安全属性,Exposure标记失败变得可观察之处,Response评估遏制、修复或撤销。我们的综述揭示,失败不一定源于有害信息:当适应将其持久性、权威性或范围扩展到其有效条件之外时,合法的状态可能变得不安全。现有工作对准入、检索、激活、暴露和局部遏制提供了相对较强的证据,但对后代修复和适应恢复后的评估提供的证据要少得多。因此,我们主张进行纵向评估,追踪不安全影响到其起源转变,验证跨后代的修复,并测试其在持续进化下是否可能重新出现。

英文摘要

Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, memories, tool definitions, skills, and workflows-from data, feedback, and accumulated experience. This shift changes the safety problem: once experience becomes reusable state, past events become future causes, and information harmless in one context may later influence decisions with greater persistence, authority, or scope. Self-evolving agent safety therefore asks not only whether a response is aligned or an action authorized, but whether safety properties survive the accumulation, generalization, and cross-context reuse of locally useful experience. We introduce SAVER, a transition-centered framework in which Substrate locates reusable influence, Adaptation captures how it changes, Violation identifies compromised safety attributes, Exposure marks where failures become observable, and Response assesses containment, repair, or revocation. Our survey reveals that failures need not originate from harmful information: legitimate state can become unsafe when adaptation expands its persistence, authority, or scope beyond the conditions under which it was valid. Existing work provides comparatively strong evidence for admission, retrieval, activation, exposure, and local containment, but much less for descendant repair and evaluation after adaptation resumes. We therefore argue for longitudinal evaluation that traces unsafe influence to its originating transition, verifies repair across descendants, and tests whether it can re-emerge under continued evolution.

CommentsSurvey paper; 80 pages, 6 figures, 13 tables. Project page: https://xaddwell.github.io/Awesome-Self-Evolving-Agent-Safety/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑