arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01436cs.AI

一个具有ATLAS对齐可执行规则和形式验证的确定性与可审计AI安全风险评估框架

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

Yixuan Huang, Basel Halak, Boojoong Kang

首次发表
浏览论文内容

中文总结 AI 辅助

该框架提出一种确定性、可审计的AI安全评估方法,通过ATLAS对齐的可执行规则和形式验证,实现从工程工件到技术级结果的机器可评估映射,并在开源项目中验证有效性。

中文摘要 AI 辅助

人工智能系统越来越多地部署在高影响力和安全关键的场景中,然而安全评估在审计下仍然难以复现和辩护。现有方法通常依赖于叙述性检查表或评估者驱动的评分,并且缺乏从可观察的工程工件到稳定的技术级结果的显式、机器可评估映射。我们提出了一种证据驱动的AI安全评估框架,将评估操作化为确定性决策函数。该框架将异构工件规范化为项目无关的控制ID分类法,在有限四级有序量表上评分,通过显式的缓解到控制映射从固定的MITRE ATLAS快照编译技术级谓词,并输出技术索引的可行性和影响级别,以及追溯到触发证据的可追踪链接。我们将所有规范性选择打包为版本化评估策略对象,以支持跨快照的可重复重新评估。为确保语义正确性,我们在完整声明的评分域上形式验证编译评估器的有界性、完全性、有序语义一致性和单调性。我们在五个固定到显式仓库快照的公共开源AI项目上评估该框架,量化统一加固干预下变更前后的差异,并通过基于分叉的软件物料清单(SBOM)生成和持续集成(CI)安全扫描门实现验证对真实工程变更的响应性。结果表明,在增强的可观察控制下,可行性概况持续向下移动,而当技术特定核心控制仍不在证据范围内时,最坏情况残余可行性仍然存在。

英文摘要

Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.

发表机构

  • University of Southampton(南安普顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑