面向软件交付决策关卡的智能体安全审计器持续保证
Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates
浏览论文内容
中文总结 AI 辅助
针对LLM代码库审计器的非确定性导致时间点审计无法为软件合并决策提供持续保证的问题,提出政策-证据-执行分离模式及TAIP保证引擎,实验验证其在多上下文下的重计算延迟远低于预算要求。
中文摘要 AI 辅助
基于大语言模型(LLM)的代码库审计器正越来越多地作为安全控制措施部署在持续集成(CI)流水线中,其审计结果会决定软件变更是被准入、阻断还是延迟。作为智能体软件开发生命周期(SDLC)安全控制,它们的非确定性行为会改变证据,而组织的风险偏好以及司法管辖或数据主权政策会改变对证据的解读。因此,时间点审计无法为合并决策维持当前的保证状态。\n我们提出了“政策-证据-执行分离模式”,由可信AI态势(TAIP)保证引擎实现,并以持续控制态势保证(CCPA)模式运行。通过将政策与稳定的执行分离,并将准入证据绑定到带版本控制的态势树(Posture Tree),相同的保证逻辑可跨模型、环境和政策配置运行。\n我们使用未修改的RepoAudit在固定的Python空指针解引用基准测试上评估了该方法。留存的证据库包含跨两种OpenAI模型配置(gpt-4o-mini和gpt-4.1)的80次RepoAudit执行。TAIP会在政策、证据和模型上下文发生变化后重新计算保证态势,并在数量不断增加的独立决策网关(Decision Gateway)上下文中进行评估。在一次执行的三个政策类周期中,观测到的政策到态势的最大延迟为1.1毫秒。在1000个独立保证上下文下,单工作进程的全政策触发重计算记录的最大总刷新时间为1.62秒,低于预先声明的5秒决策网关预算。这些单主机测量针对的是留存证据的保证过程,不包含RepoAudit执行和服务商推理的时间。
英文摘要
Large language model (LLM)-based repository auditors are increasingly deployed as security controls within continuous integration (CI) pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle (SDLC) Security Controls, their non-deterministic behaviour changes the evidence, while organisational risk appetite and jurisdictional or data-sovereignty policy change its interpretation. Point-in-time audits therefore cannot maintain current assurance for merge decisions. We propose the Policy-Evidence-Execution Separation Pattern, implemented by the Trustworthy AI Posture (TAIP) Assurance Engine and operated as Continuous Control Posture Assurance (CCPA). By separating policy from stable execution and binding admitted evidence to a versioned Posture Tree, the same assurance logic operates across models, environments, and policy profiles. We evaluate the approach using unmodified RepoAudit on a fixed Python Null Pointer Dereference benchmark. The retained evidence repository contains 80 RepoAudit executions across two OpenAI model configurations, gpt-4o-mini and gpt-4.1. TAIP recomputes assurance posture after policy, evidence, and model-context changes and is evaluated across increasing numbers of independent Decision Gateway contexts. The maximum observed policy-to-posture latency was 1.1 ms across three policy-class cycles in one execution. At 1,000 independent assurance contexts, full policy-triggered recomputation with one worker recorded a maximum aggregate refresh of 1.62 s, below the predeclared 5 s Decision Gateway budget. These single-host measurements concern assurance over retained evidence and exclude RepoAudit execution and provider inference.
发表机构
- Swinburne University of Technology(斯威本科技大学)
- CSIRO(澳大利亚联邦科学与工业研究组织)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。