arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36580cs.AI

SafeCoEvo:测试时协同进化LLM智能体的安全护栏与守卫

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin, Weicheng Meng, Jingyang Qiao, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Weilin Luo, Kun Shao, Dong Li, Zhizhong Zh… 展开作者

Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin, Weicheng Meng, Jingyang Qiao, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Weilin Luo, Kun Shao, Dong Li, Zhizhong Zhang, Yuan Xie, Zhaoxia Yin

首次发表
浏览论文内容

中文总结 AI 辅助

SafeCoEvo提出测试时护栏与守卫协同进化框架,通过短期显式知识更新和长期参数化风险判断,同时降低不安全结果率10.05%并提升任务成功率12.15%。

中文摘要 AI 辅助

部署在真实环境中的LLM智能体会持续遇到新任务和安全风险,而执行反馈通常仅在每个任务完成后才可用。然而,现有的自进化方法通常依赖于在固定且可重复访问的任务分布上进行多轮优化,这与真实部署中的测试时自适应根本不同,因为在真实部署中,只能利用过去任务积累的经验来改善未来未见任务上的安全决策。为解决这一局限,我们提出了SafeCoEvo,一个用于LLM智能体安全的测试时护栏-守卫(Harness-Guard)协同进化框架,使外部安全系统能够从积累的运行时经验中持续适应。SafeCoEvo在不同时间尺度上联合改进两种互补的安全能力:S-Harness将近期运行时经验快速外部化为可更新的显式安全知识,从而能及时影响后续任务;而GuardVPO在更长的时间尺度上将积累的运行时安全经验内化为参数化的风险判断能力。通过将短期快速适应与长期能力巩固相结合,SafeCoEvo持续提升智能体的安全能力,在最强基线上将不安全结果率降低了10.05%,同时将任务成功率提高了12.15%,从而在安全性和任务效用上实现同步提升。

英文摘要

LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.

发表机构

  • East China Normal University(华东师范大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Xiamen University(厦门大学)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • University College London(伦敦大学学院)
  • Huawei Noah’s Ark Lab, UK(华为诺亚方舟实验室(英国))
  • MemoraX AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑