面向军事指挥与控制的智能体AI系统的测试与评估
Testing and Evaluation of Agentic AI Systems In Military Command and Control
浏览论文内容
中文总结 AI 辅助
本文针对军事C2场景中智能体AI系统测试与评估的保障问题,分析既定方法的假设被智能体特性削弱的情况,推导保障主张并评估现有方法,明确部分举证责任转移至部署阶段。
中文摘要 AI 辅助
公共承诺要求对用于军事指挥与控制(C2)的智能体AI系统进行严格测试并实施人工监督,而这些承诺能否兑现取决于支撑其的保障论证,该论证需包含三个要素:规定可接受条件的主张、与这些主张相关的证据,以及连接二者的论证。通过对240项已记录的测试与评估(T&E)实践进行结构化审查,涵盖8个评估维度和3个生命周期阶段,我们确定了既定方法对其测试对象作出的8项假设,这些假设归为4个集群:系统可指定性、稳定性、可组合性和可监督性。智能体特性会削弱所有8项假设,这种削弱影响的是连接证据与主张的论证,而非主张或证据本身。因此,测试结果可能满足流程要求,但无法为从测试行为到部署后行为的推断提供保障。我们针对前三个假设集群推导了10项保障主张,并评估当前及新兴方法能否应对每项主张,通过5个C2场景映射了作战后果。可监督性已被确定,但在此未予评估,因为证明其成立需依赖系统稳定性结果及超出本研究范围的人因T&E方法。已记录的现有证据不支持关于系统级行为的宽泛主张,但较窄范围的主张原则上仍可恢复,前提是具备成熟方法:受限任务包络、基于轨迹的正确性、可执行运行时约束,以及已表征的运行间方差。部分举证责任转移至部署阶段,这使得部署决策成为一个持续的行为。若无法生成证据,可通过定义明确的失效条件和指定责任方来管控剩余不确定性。
英文摘要
Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable in principle, contingent on mature methods: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.
发表机构
- Arcadia Impact(阿卡迪亚影响机构)
- AI Governance Taskforce(AI治理工作组)
- Veraitech(维瑞泰克公司)
- University of Oxford(牛津大学)
- King’s College London(伦敦国王学院)
- Atlantic Council(大西洋理事会)
- Future Ethics Lab(未来伦理实验室)
机构由 AI 辅助整理,请以论文原文为准。