幻影护栏:当自我改进的智能体利用从未发生过的修复失败时
Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened
浏览论文内容
中文总结 AI 辅助
研究自我改进智能体在自动化工具优化中虚构从未发生的错误这一现象,通过反事实制造实验室进行研究,发现当合法输入含特定模式等条件时会出现虚构失败,提出该实验室用于测量此类虚构失败。
中文摘要 AI 辅助
自我改进的人工智能智能体旨在从错误中学习。我们发现它们也会虚构从未发生过的错误。我们在自动化工具优化中研究这种失败模式,基于大语言模型的提议者编辑智能体的框架,包括提示、解析器、过滤器、验证器和护栏,以消除观察到的失败。但这个过程很少先问:是否真的有失败需要修复?我们引入了反事实制造实验室,这是一个确定性的微实验室,其中正确的行动是已知的:什么都不做。该实验室为一个可证明从未发生过的失败类植入一个候选护栏,只呈现合法情节,并使用字节精确的预言机来检查每一个引用的违规行为。提议者在实际违规行为上按预期行事,对无特征的合法输入弃权。然而,当合法输入包含一个类似于熟悉游戏规则的无害模式时,它会虚构一个失败:在60次运行中有15次,而在无特征输入上为0/60,它启用不存在规则的护栏并引用预言机反驳的违规行为。这种影响是有结构的,而不是随意的。在单次提议中,只有当三个条件同时出现时才会出现:规则形状的模式、开放式规则集和预设失败的指令。去除任何一个条件都能消除制造。因为虚构的护栏不会改变任何真实结果,也无法提高已经完美的抑制分数,这种现象既不是奖励黑客行为也不是过度拒绝。它是一个幻影护栏:对从未发生过的失败的修复,对于仅抑制接受是不可见的。在仅添加接受循环中,即使没有预设失败的指令,它也会重新出现,循环的持续添加角色提供了单次提议中指令提供的需求,一旦进入就会停留。我们提出了反事实制造实验室来测量自我改进智能体工具中的虚构失败。
英文摘要
Self-improving AI agents are designed to learn from their mistakes. We show they can also hallucinate mistakes that never happened. We study this failure mode in automated harness optimization, where an LLM-based proposer edits an agent's scaffold, including prompts, parsers, filters, validators and guardrails, to eliminate observed failures. But this process rarely asks first: was there a real failure to fix? We introduce the Counterfactual Fabrication Lab, a deterministic micro-lab where the correct action is known: do nothing. The lab plants a candidate guardrail for a failure class that provably never occurs, presents only legal episodes, and uses a byte-exact oracle to check every cited violation. The proposer behaves as expected on real violations and abstains on featureless legal input. Yet when the legal input contains a harmless pattern resembling a familiar game rule, it invents a failure: in 15/60 runs, versus 0/60 on featureless input, it enables the nonexistent-rule guardrail and cites a violation the oracle refutes. The effect is structured, not indiscriminate. In single-shot proposals it appears only when three conditions coincide: a rule-shaped pattern, an open-ended rule set and an instruction that presupposes failures. Removing any of these conditions eliminates the fabrication. Because the invented guardrail changes no true outcome and cannot improve an already-perfect suppression score, the phenomenon is neither reward hacking nor over-refusal. It is a phantom guardrail: a fix for a failure that never happened, invisible to suppression-only acceptance. Inside an add-only accept loop it re-enters even without the failure-presupposing instruction, the loop's keep-adding role supplying the demand the instruction supplied in single shot, and once in it stays. We present the Counterfactual Fabrication Lab for measuring fabricated failures in self-improving agent harnesses.