arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08175cs.AI

安全框架自进化:可行性与局限性的理论分析

A Theory of Reliable Self-Evolution for Agent Harnesses

发表机构中国科学技术大学 · 香港生成式人工智能研发中心 · 香港科技大学
另 1 家 · 查看机构详情
  • University of Science and Technology of China(中国科学技术大学)
  • Hong Kong Generative AI Research & Development Center(香港生成式人工智能研发中心)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Hong Kong Baptist University(香港浸会大学)

机构由 AI 辅助整理,请以论文原文为准。

Qianshu Cai, Yonggang Zhang, Jun Nie, Maohao Ran, Huajiang Zheng, Jun Song, Xinmei Tian, Yike Guo, Wei Xue

首次发表
浏览论文内容

中文总结 AI 辅助

本文从理论上分析了安全框架自进化的可行性与局限,揭示了生成与认证的约束差异及停滞原因,为设计更安全的自我进化机制提供了基础。

中文摘要 AI 辅助

框架自进化是指智能体根据任务反馈修改其提示、工具、代码或编排,同时保持底层语言模型冻结,且修改在后续任务中持续存在的过程。我们针对安全框架自进化的可行性与局限性进行了系统的理论分析,将修改生成、有限数据认证与选择、安全采纳以及更新后的行为联系起来。在固定的用户任务分布下,我们建立了保证整体期望奖励提升同时控制保留任务上变化的条件,刻画了生成合格修改的概率,并推导了安全选择与采纳的有限数据界。我们的分析表明,生成与认证施加了不同的约束:当前任务性能并不能决定生成合格修改的概率,且在评估受限时,生成更多候选并不必然提高成功更新的保证。因此,即使改进机会仍然存在,也可能出现停滞。我们进一步证明,当期望奖励接近其上界时,识别真正改进的最坏情况评估成本会发散。在连续更新过程中,认证的改进保证在有限运行内累积,但成功更新本身并不保证进一步改进仍然可能。这些结果为诊断瓶颈和设计更安全的自我进化机制提供了基础。

英文摘要

In harness self-evolution, agents modify their own prompts, code, tools, and orchestration while keeping the underlying language model fixed. Recent work has shown that agents can improve themselves in response to task failures and achieve substantial performance gains. However, gains on failed tasks do not automatically ensure that performance on previously successful tasks is preserved, raising concerns about reliable adoption. In this work, we provide a theoretically grounded condition under which a self-evolved harness can be reliably adopted. We then propose a validation rule to make the evolved system satisfy the reliable adoption condition with theoretical guarantees. Consequently, the system can achieve progressive improvement through evolution. This leads to a natural question: Does reliable self-evolution have a performance ceiling, and which factors govern this ceiling? Our theoretical results show that the performance ceiling is determined by the costs of verification and evaluation. On the other hand, self-evolution may stall in practice. In this regard, we show that experimental evidence on agent performance shifts can be used to identify the sources of stagnation. Our work thus establishes a theoretical framework for understanding and advancing reliable harness self-evolution.

↑