模型提议,代码处置:一个由LLM编排的进攻性安全智能体中验证器与接受阶段的预注册消融研究
The Model Proposes, the Code Disposes: A Pre-Registered Ablation of a Verifier-and-Acceptance Stage in an LLM-Orchestrated Offensive-Security Agent
浏览论文内容
中文总结 AI 辅助
本研究通过预注册消融实验证明,在LLM编排的进攻性安全智能体中,由确定性代码强制执行的模型验证器阶段能显著抑制虚假报告并提高交付精确率,但未确立对实验室目标之外的泛化优势。
中文摘要 AI 辅助
我们评估了一个验证器与接受阶段——即由确定性代码强制执行其裁决的模型验证器——是否会改变LLM驱动的进攻性安全智能体所报告的内容。我们报告了一项15次运行的探索性试点、一项预注册的20次运行确认性消融研究,以及一项预注册的2×2因子设计研究,该研究在两个故意存在漏洞的实验室目标上进行了40次运行。在确认性研究中,移除该阶段消除了报告前抑制(每次运行的中位数从2降至0;精确单侧p=0.00003),并降低了模型盲发的已交付精确率(中位数从0.471降至0.353;p=0.0087)。针对一个冻结但不完整的地面真值列表的召回率没有显著差异(双侧p=0.158;未建立等价性)。因子设计研究将抑制归因于模型验证器(Holm调整后p=0.004);仅确定性接受规则未抑制任何假阳性,且未检测到交互作用(p=0.72)。完整设计保留了模型裁决的真候选者的93.8%,但未达到其预注册的非劣效性标准,因为单侧95%置信下限为0.875,低于0.90的下限。在确认性和因子设计研究中,一个仪表化的金丝雀在60次运行中的60次中记录了零接触,偶发的外部接触已单独披露。对保留的盲包的独立人工裁决尚未完成,因此精确率和敏感性终点是支持性而非最终证据。还披露了六次审计跟踪失败,包括一次发生在评估工具中。结果支持一个狭窄的结论:验证器改变了系统交付的内容,而确定性代码提供了执行和可审计性;这些结果并未确立相对于其他智能体的优越性或超越实验室目标的泛化能力。
英文摘要
We evaluate whether a verifier-and-acceptance stage - a model verifier whose verdicts are enforced by deterministic code - changes what an LLM-driven offensive-security agent reports. We report a 15-run exploratory pilot, a pre-registered 20-run confirmatory ablation, and a pre-registered 2 x 2 factorial study with 40 runs across two deliberately vulnerable lab targets. In the confirmatory study, removing the stage eliminated pre-report suppression (median 2 versus 0 findings per run; exact one-sided p = 0.00003) and reduced model-blinded shipped precision (median 0.471 versus 0.353; p = 0.0087). Recall against a frozen but incomplete ground-truth list did not differ significantly (two-sided p = 0.158; equivalence was not established). The factorial study attributed suppression to the model verifier (Holm-adjusted p = 0.004); deterministic acceptance rules alone suppressed no false positives, and no interaction was detected (p = 0.72). The full design retained 93.8% of model-adjudicated true candidates but did not meet its pre-registered non-inferiority criterion because the lower one-sided 95% bound was 0.875, below the 0.90 floor. Across the confirmatory and factorial studies, an instrumented canary recorded zero contacts in 60 of 60 runs, with incidental external contacts disclosed separately. Independent human adjudication of the retained blind packets is pending, so precision and sensitivity endpoints are supporting rather than final evidence. Six audit-trail failures, including one in the evaluation tooling, are also disclosed. The results support a narrow conclusion: the verifier changes what the system ships, while deterministic code supplies enforcement and auditability; they do not establish superiority to other agents or generalization beyond lab targets.