arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

权限受限:强化环境下编码智能体的策略分级评估

Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments

Dotan Davidovich, Yair Amar, Hai Rozencwajg, Or Hiltch

arXiv 2608.02670首次发表:更新:

发表机构

Accomplish AI(Accomplish AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对强化环境下编码智能体的性能评估问题,在 Terminal-Bench 2.1 上评估 12 个智能体,发现策略强化会导致性能损失且模型选择依赖策略,并发布 Boundary-Bench 插件支持相关评估。

AI 中文摘要

编码智能体越来越多地在组织内部运行,这些组织的安全控制措施(如范围受限的凭证、受限的出口访问、只读文件系统、非根执行)会像约束其他任何软件一样对其进行限制。然而,现有的基准测试几乎仅在宽松的沙箱环境中评估智能体,因此当实施严格策略时性能会如何变化尚不明确。在本研究中,我们在 Terminal-Bench 2.1 上评估了 12 个编码智能体,这些智能体面临源自常见真实企业限制的嵌套安全策略级别。强化措施并非无代价,但其代价差异极大:在最严格的策略下,成功率损失达 18.3 个百分点,成本通胀达 167.3%,且这两个指标并不一致;最能保持成功率的模型也是效率损失最大的模型,因此模型选择取决于具体策略。除聚合分数外,我们还描述了智能体在策略阻止其操作时的行为,并分解了强化措施引发的失败:运行过程会陷入超时或错误解决方案,而非提前终止,且这种情况的组合因模型而异。为了为比较提供依据,我们验证了最严格策略下的任务可解决性,将模型失败与策略禁止的任务区分开来。我们发布了 Boundary-Bench,这是一个开源的强化插件,可支持在 Terminal-Bench 及兼容基准上对编码智能体进行策略受限评估。

英文摘要

Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root execution) constrain them like any other software. Existing benchmarks, however, evaluate agents almost exclusively in permissive sandboxes, so it is unknown how performance changes when policy is enforced. In this work, we evaluate 12 coding agents on Terminal-Bench 2.1 across nested security policy levels derived from common real-world enterprise restrictions. Hardening is never free but far from uniform: under the strictest policy, success losses reach 18.3 points and cost inflation 167.3\%, and the two axes disagree; the model that best preserves success is also the one that loses the most efficiency, so model choice is policy-dependent. Beyond aggregate scores, we characterize how agents behave when policy blocks their actions and decompose the failures hardening induces: runs grind into timeouts or wrong solutions rather than stopping early, in a mix that differs by model. To ground comparisons, we verify task solvability under the strictest policy, separating model failures from tasks the policy forecloses. We release Boundary-Bench, an open-source hardening plugin enabling policy-constrained evaluation of coding agents on Terminal-Bench and compatible benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑