发表机构
Tulane University; University of British Columbia; Rutgers University; Princeton University; New York University; Virginia Tech(杜兰大学; 不列颠哥伦比亚大学; 罗格斯大学; 普林斯顿大学; 纽约大学; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过有状态授权任务的测试平台,探讨递归自我改进中安全失败的持续原因,并提出通过刷新分数和编辑初始实现等干预措施来改善恢复,同时保留部署效率。
AI 中文摘要
递归自我改进(RSI)使智能体能够将有用的变更跨代传递。在跨代维护安全涉及两个方面:防止不安全行为的持续存在,以及在失败发生时实现恢复。我们通过一个有状态授权任务的受控测试平台研究这些挑战,其中固定的LLM编辑器优化可执行的智能体组件,独立的轨迹记录它们的效果。配对干预区分哪些修订通过验证、哪个程序继续运行以及编辑器接下来修订哪个程序。当新的授权依赖使先前测试的优化失效时,历史分数在48个框架历史中的22个中保留了相同的不安全程序,尽管每个受影响的存档中都有正确的替代方案。刷新分数在原始套件上恢复了正确性,但在独立组合的测试上仍有残余失败。在不变契约下,当所有提案被拒绝且失败的现任者保持活跃时,失败也会持续存在。从共同的失败开始,编辑初始实现而非失败的那个可改善恢复,尽管优势因编辑器而异。完整验证和经过验证的回滚在核心轨迹研究中以完全正确的程序结束,同时保留超过43%的部署节省。维护智能体安全需要检查在当前条件下将运行什么,并选择接下来编辑哪个实现。
英文摘要
Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization tasks, where fixed LLM editors optimize executable agent components and independent traces record their effects. Paired interventions separate which revisions pass validation, which program continues running, and which program the editor revises next. After a new authorization dependency invalidates previously tested optimizations, historical scores preserve the same unsafe programs in 22 of 48 framework histories despite a correct alternative in every affected archive. Refreshing scores restores correctness on the original suite, with residual failures on independently composed tests. Failures also persist under an unchanged contract when all proposals are rejected and the failed incumbent remains active. Starting from shared failures, editing the initial implementation instead of the failed one improves recovery, although the advantage varies across editors. Full validation and validated rollback end with fully correct programs in the core trajectory study while retaining over 43% deployment savings. Preserving agent safety requires checking what will run under current conditions and choosing which implementation to edit next.
Comments44 pages, 13 figures. Project page: https://RSI-Safety.github.io/