arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使智能体部署安全民主化:一种结构监测方法

Democratizing Agent Deployment Safety: A Structural Monitoring Approach

Preeti Ravindra, Rahul Tiwari, Vincent Wolowski

arXiv 2607.14570首次发表:更新:

AI 中文总结

研究人工智能智能体部署安全问题,引入信息流图(IFG)监测器,通过控制流、数据流图及原始代码差异分析结构安全回归。实验表明,IFG监测器能降低攻击错过率,还可同步作为部署前保障,为组织实现部署安全民主化提供实用途径。

AI 中文摘要

人工智能软件开发智能体越来越有能力修改基础设施和安全关键系统,在完成任务时可能暗中削弱安全保障。许多采用编码智能体的组织和个人用户缺乏维护复杂监测组件的资源和治理成熟度。本文将此问题视为基础设施即代码设置中任务成功情况下的基础设施层面破坏行为,比较同步和异步部署模式下训练和未训练的监测器。引入信息流图(IFG)监测器,通过控制流、数据流图差异及原始代码差异分析结构安全回归。异步评估中,未训练的git差异监测器在1%误报率审计预算下错过11.6%的攻击,未训练的IFG监测器降至3.5%,训练后的异步随机森林监测器低于1%。IFG还可同步作为部署前保障,在同步模式下,IFG实时回滚将联合主任务和隐蔽任务成功率从74.4%降至0.0%,且合法任务完成率无显著降低。这些结果表明,未训练的结构监测器为组织实现先进人工智能智能体的可信采用提供了一条实用且可审计的部署安全民主化途径。

英文摘要

AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms. While frontier laboratories may deploy sophisticated monitoring pipelines, many organizations and individual users adopting coding agents lack the resources and governance maturity required to maintain complex learned monitor ensembles. We study this problem as infrastructure level sabotage under task success in the infrastructure-as-code setting from ControlArena and compare trained and untrained monitors across synchronous and asynchronous deployment modes. We introduce an Information Flow Graph (IFG) monitor that analyzes structural security regressions using control-flow and data-flow graph diffs alongside raw code diffs. In asynchronous evaluation, an untrained git diff monitor misses 11.6% attacks at 1% false positive rate auditing budget, our untrained IFG monitor reduces this to 3.5%, and the trained Async RF monitor achieves below 1%. Unlike the trained Async RF monitor, IFG can also operate synchronously as a pre-deployment safeguard that blocks suspicious infrastructure changes before execution or merge. In synchronous mode, IFG realtime rollback reduces joint main task and covert task success from 74.4% to 0.0% with no measurable reduction in legitimate task completion. These results suggest that untrained structural monitors provide a practical and auditable path toward democratizing deployment safety for organizations seeking trustworthy adoption of advanced AI agents.

CommentsAccepted in ICML 2026 Workshops: AI4GOOD and AIWILD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑