智能体自修改运行时监督的自愈防护装置
Self-Healing Harness for Runtime Oversight of Agent Self-Modification
浏览论文内容
中文总结 AI 辅助
针对LLM智能体自修改的持久化控制问题,提出外部运行时门控的自愈防护装置,通过检测-通知-修复-验证循环实施准入控制,实验表明能有效拦截带附带回归的修改并提升任务完成度。
中文摘要 AI 辅助
LLM智能体能够改变自身未来的行为,这引发了一个基本的控制问题:哪些自生成的更改应被允许持久化。我们将其表述为自修改的准入控制。智能体可以对其操作指令提出更改,而外部运行时门控则控制持久性。我们将这一原则实现为一个与模型无关的自愈防护装置,围绕一个未经修改的智能体运行“检测、通知、修复、验证”循环。智能体在外部工作区中编写候选行为规则,在评估期间获得临时执行权限,并且仅在触发失败上测得改进且受保护案例的回归不超过固定裕度后,才获得持久的跨回合权限。重放提供可用的匹配证据,前向试验提供较弱的回退方案,语料库级守卫重新测试累积的活动规则集。在跨越AppWorld、Terminal-Bench和τ²-Bench的16组匹配的基线与防护装置运行中,门控拒绝了383个由重放决定的提案。其中,211个(55%)在改进其触发失败的同时,使先前正常工作的案例退化。这表明局部有益的自修改可能足够频繁地引入附带回归,从而实质性影响门控决策,为外部准入控制提供了直接的经验动机。在所有16对中,防护装置下的任务完成分数更高,其中两对的配对自助法区间不包含零,而重复试验可靠性在12对中更高,4对持平,无一更低。由于适应修改了策略诱导上下文而保持模型权重不变,被接受的更改保持可检查、可逆,并与闭权重模型兼容。
英文摘要
LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an external runtime gate controls persistence. We implement this principle as a model-agnostic self-healing harness that runs a Detect, Notice, Heal, Validate loop around an otherwise unmodified agent. The agent authors candidate behavioral rules in an external workspace, where they receive provisional execution authority during evaluation and acquire persistent cross-episode authority only after measured improvement on the triggering failure without regression beyond a fixed margin on protected cases. Replay provides matched evidence when available, forward trials provide a weaker fallback, and a corpus-level guard re-tests the accumulated active rule set. Across 16 matched Baseline and Harness runs spanning AppWorld, Terminal-Bench, and $τ^2$-Bench, the gate rejected 383 replay-decided proposals. Of these, 211 (55%) improved their triggering failure while degrading a case that previously worked. This shows that locally beneficial self-modifications can introduce collateral regressions often enough to materially affect gate decisions, providing direct empirical motivation for external admission control. Task-completion score is higher under the Harness in all 16 pairs, with two paired bootstrap intervals excluding zero, while repeated-trial reliability is higher in 12 pairs, tied in 4, and lower in none. Because adaptation modifies the policy-inducing context while leaving model weights fixed, admitted changes remain inspectable, reversible, and compatible with closed-weight models.
发表机构
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- AI Labs, Capital One(第一资本人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。