arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

构建者、防御者、破坏者:反对从人工智能驱动的安全生命周期中移除人类的理由

Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test Software

Mohamed Chahine Ghanem

arXiv 2607.03215首次发表:更新:

发表机构

School of Computer Science and Mathematics, Keele University; Cybersecurity Institute, University of Liverpool(基尔大学计算机科学与数学学院; 利物浦大学网络安全研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

探讨人工智能在安全生命周期中集多角色于一身的现象,指出移除人类会带来问题,基于多领域证据论证人类应是安全生命周期的永久结构要求,并阐述人机合理分工。

AI 中文摘要

人工智能已贯穿整个安全生命周期,同一类模型同时承担编写、强化和探测应用代码的角色。本文认为,构建工件的系统与防御和测试它的系统若来自同一分布,会导致问题,移除人类会引发诸多弊端,人类应是安全生命周期的永久结构要求并给出人机分工建议。

英文摘要

Generative models now write application code, harden and monitor it, and probe it for exploitable flaws, so that one family of models increasingly plays builder, defender and breaker at once. The prevailing view treats full autonomy as the natural end point of assistance. This article argues for a narrower and more defensible position than a blanket requirement for human oversight. We define the shared generative substrate as the set of upstream dependencies (training corpus, model family, alignment procedure, vendor, toolchain) that induce correlated errors, and we define independence as a measurable property of pairs of lifecycle roles rather than an attribute of any participant. A coincident-failure model in the tradition of Eckhardt and Lee shows why organisational independence, the proxy on which verification standards have relied, no longer implies statistical independence once roles share a substrate; it also shows that heterogeneous models, deterministic analysers and formal verification can restore independence for the fault classes they cover. What remains for humans is specific: authority over the specification against which every machine oracle is judged, accountability that current governance instruments do not permit to be delegated, and last-resort authority to halt. We translate this into an operational framework with five autonomy levels, three distinct human roles and five decision criteria (consequence, reversibility, time-criticality, verifiability and adversarial exposure), apply it to build, defend and test actions, and specify how independence and oversight effectiveness can be measured. The central claims are stated as testable hypotheses with a protocol for refuting them.

Comments21 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑