arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

治理即代码:将欧盟人工智能法案技术要求转化为生成式人工智能系统的可执行合规流水线

Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems

Rudrendu Kumar Paul, Sourav Nandy

arXiv 2609.20016首次发表:更新:

发表机构

Boston University; University of Texas at Austin(波士顿大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对欧盟人工智能法案在生成式系统上的技术缺口,提出治理即代码框架,以43项可检查标准在CI/CD中自动审计,验证并削减75%审计工作量。

AI 中文摘要

欧盟人工智能法案(第2024/1689号法规)对高风险人工智能提供者施加了技术义务,然而第8至15条是为预测性人工智能起草的,在应用于生成式系统时留下了七个技术缺口,涵盖非确定性数据治理、训练数据来源、持续合规、人类监督、开放式鲁棒性、新兴风险以及生成式公平性。我们提出了治理即代码(GaC)框架,该框架包含六个合规模块中的43项可机器检查的验收标准,这些标准在CI/CD流水线中运行并生成按条款索引的审计证据,我们展示了实际的Rego策略代码而非仅作描述。我们的核心承诺是,法案中开放性文本标准(如“适当水平”、“可能的偏见”)转化为已声明且可审计的数字:鲁棒性阈值源自提供者记录在案的基线和最先进水平的下限,框架偏见被分解为八个可测量的代理指标,并通过反事实人口统计探测进行测试。我们还纠正了谁应承担何种责任,因为根据第25条和第五章,下游部署者依赖上游提供者根据第53条提供的训练数据摘要,仅记录其控制的层,因此GaC验证该摘要而非要求部署者从未拥有的逐样本文档。我们在两个企业部署中进行了验证,一个是高风险咨询聊天机器人,另一个是有限风险内容生成器,与手动专家审计而非从未设计用于执行合规的文档工件进行基准比较。GaC复现了手动审计的所有发现,包括三项触发处罚的违规行为,同时将审计工作量削减约75%。

英文摘要

The EU AI Act (Regulation 2024/1689) imposes technical obligations on high-risk AI providers, yet Articles 8-15 were drafted for predictive AI and leave seven technical gaps when applied to generative systems, spanning non-deterministic data governance, training-data provenance, continuous conformity, human oversight, open-ended robustness, emergent risk, and generative fairness. We deliver Governance-as-Code (GaC), a framework of 43 machine-checkable acceptance criteria across six compliance modules that run in a CI/CD pipeline and emit Article-indexed audit evidence, and we show the actual Rego policy code rather than merely describing it. Our central commitment is that the Act's open-textured standards ("appropriate levels," "possible biases") become declared, auditable numbers: robustness thresholds are derived from the provider's documented baseline and a state-of-the-art floor, and framing bias is collapsed into eight measurable proxies tested by counterfactual demographic probing. We also correct who owes what, since under Article 25 and Chapter V a downstream deployer relies on the upstream provider's Article 53 training-data summary and documents only the layers it controls, so GaC verifies that summary rather than demanding per-sample documentation the deployer never had. We validate on two enterprise deployments, a high-risk advisory chatbot and a limited-risk content generator, benchmarking against a manual expert audit rather than documentation artifacts that were never designed to enforce compliance. GaC reproduces all of the manual audit's findings, including three penalty-triggering violations, while cutting audit labor by roughly 75%.

CommentsAccepted at the AI4Law Workshop, ICML 2026. Camera-ready version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑