arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24200cs.CR

自动化计算机安全测试中的可伪造确认:确定性规则与AI裁判

Forgeable Confirmation in Automated Computer Security Testing: Deterministic Rules versus AI Judges

  • OmiCore Inc.(OmiCore公司)
  • Graduate School of Engineering, Kyushu University(九州大学研究生院)

机构由 AI 辅助整理,请以论文原文为准。

Akihisha Fujiyama, Niwase Shamim

AI总结:

本研究证明在AI辅助安全测试中,读取攻击者控制数据的确认机制可被伪造,并提出将决定性证据移至不可写通道以将攻击成功率从97%降至0%。

AI中文摘要:

人工智能越来越多地被用于自动化计算机安全测试,这些工具必须自行判断攻击是否成功。由确定性规则通过观察确认的发现被报告为事实,而由大语言模型(LLM)判定为可利用的发现则被视为意见。我们探究被测系统能否伪造这种确认。在对一个四阶段AI辅助管道的离线安全测试中,其十五种确认机制中有九种是可伪造的,而可伪造性完全取决于决策是否读取攻击者控制的数据。我们将此形式化为一个可审计的攻击面,并进行前瞻性测试:在十六种保留的机制上,攻击前确定的预测精确区分了可伪造与不可伪造的机制,而在公开扫描器模板中的12,203种机制上,预测准确率达99.9%。确定性规则比八个开放权重LLM裁判更易被伪造,在攻击者控制的响应内容上失败率为2%,而LLM裁判的中位失败率为50%。没有一种检查的实现既稳健又精确,而在规则与AI裁判之间进行路由将伪造成功率提升至99%。将决定性证据移至攻击者无法写入的通道,可将攻击成功率从97%降至0%,而升级(escalate)判定可恢复由此损失的敏感性。当被扫描的主机本身是 adversary 时,该保护失效。这些结果对AI安全代理以及通过字符串匹配来评判成功率的基准测试具有意义。

英文摘要:

AI is increasingly used to automate computer security testing, and the tools must decide for themselves whether an attack succeeded. A finding that a deterministic rule confirms by observation is reported as fact, whereas one that an LLM judges exploitable is treated as an opinion. We ask whether the system under test can forge that confirmation. In offline security testing of a four-stage AI-assisted pipeline, nine of its fifteen confirmation mechanisms are forgeable, and forgeability is predicted entirely by whether the decision reads attacker-controlled data. We formalise this as an auditable attack surface and test it prospectively: on sixteen held-out mechanisms, predictions fixed before any attack separated forgeable from unforgeable mechanisms exactly, and across 12,203 mechanisms in public scanner templates the prediction was 99.9% accurate. Deterministic rules proved cheaper to forge than eight open-weight LLM judges, failing at 2% of attacker-controlled response content against a median of 50%. No implementation of one check was both robust and precise, and routing between a rule and an AI judge raised forgery to 99%. Moving the decisive evidence to a channel the attacker cannot write cuts attack success from 97% to 0%, and an escalate verdict recovers the sensitivity this costs. The protection fails when the scanned host is itself the adversary. The results bear on AI security agents and on benchmarks that score success by string matching.

补充信息

↑