arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16287cs.SE

AgentGuard:从异常编码智能体轨迹中学习执行护栏

AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories

Wuyang Dai, Song Wang

首次发表
浏览论文内容

中文总结 AI 辅助

AgentGuard 提出从编码智能体异常轨迹中自动学习指令级执行护栏,将失败模式泛化为条件约束并动态激活,实验将异常执行率从 69.0% 降至 26.7%,任务完成率从 21.7% 提升至 35.0%。

中文摘要 AI 辅助

AI 编码智能体越来越依赖执行框架(execution harnesses)与代码仓库和外部工具进行交互。然而,任务成功并不能保证执行的可靠性。智能体仍可能修改无关文件、重写测试、发出不安全命令或忽略失败的验证,这促使需要行为护栏来确保可靠执行。我们提出了 AgentGuard,一个指令级护栏框架,从编码智能体的异常轨迹中学习条件执行约束。AgentGuard 不依赖人工指定的安全规则,而是自动提取反复出现的执行失败模式,将其泛化为指令级行为约束,并组织为轻量级护栏技能,该技能仅动态激活与当前指令相关的规则。这种设计能够在实现行为引导的同时,最大限度地减少对正常执行的不必要限制。我们使用从 382 个仓库任务中收集的 642 条真实编码智能体执行失败轨迹来评估 AgentGuard。护栏从覆盖 282 个任务的 461 条轨迹中学习,并在一个不重叠的 100 个任务集上进行评估。使用 Claude Code 搭配 Claude Haiku 4.5 作为底层编码智能体,我们将基线智能体与由 AgentGuard 增强的同一智能体进行比较。实验结果表明,AgentGuard 将异常执行率(Abnormal Execution Rate)从 69.0% 降低至 26.7%,并将成功任务完成率(Successful Task Completion Rate)从 21.7% 提升至 35.0%。这些结果证明,从历史失败中学习的执行护栏能够显著提高 AI 编码智能体的可靠性,同时也凸显了在安全性与任务完成之间取得平衡的剩余挑战。

英文摘要

AI coding agents increasingly rely on execution harnesses to interact with repositories and external tools. However, task success does not guarantee reliable execution. Agents may still modify unrelated files, rewrite tests, issue unsafe commands, or ignore failed validations, motivating behavioral guardrails for reliable execution. We present AgentGuard, an instruction-level guardrail framework that learns conditional execution constraints from anomalous trajectories of coding agents. Rather than relying on manually specified safety rules, AgentGuard automatically extracts recurring execution failure patterns, generalizes them into instruction-level behavioral constraints, and organizes them as a lightweight guardrail skill that dynamically activates only the rules relevant to the current instruction. This design enables behavioral guidance while minimizing unnecessary restrictions on normal execution. We evaluate AgentGuard using 642 documented failure traces collected from real coding-agent executions across 382 repository tasks. Guardrails are learned from 461 traces covering 282 tasks and evaluated on a disjoint set of 100 tasks. Using Claude Code with Claude Haiku 4.5 as the underlying coding agent, we compare the baseline agent with the same agent augmented by AgentGuard. Experimental results show that AgentGuard reduces the Abnormal Execution Rate from 69.0% to 26.7% and increases the Successful Task Completion Rate from 21.7% to 35.0%. These results demonstrate that execution guardrails learned from historical failures can substantially improve the reliability of AI coding agents while highlighting the remaining challenge of balancing safety and task completion.

↑