发表机构
Luxembourg Institute of Science and Technology (LIST); University of Luxembourg(卢森堡科学技术研究院; 卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对欧盟AI监管沙箱中技术测试工具碎片化问题,提出开源框架AI评估沙箱配置器,通过目录化插件API、统一数据模型和角色仪表板满足11项架构治理要求,并经真实试点验证。
AI 中文摘要
欧盟《人工智能法案》要求所有成员国在2027年8月前建立AI监管沙箱(AIRS):这是一种受监管的环境,汇集了国家主管机构、技术专家以及接受评估的组织。当AIRS活动包含结构化技术测试时,大规模开展此类测试需要专用基础设施,然而工具生态系统在结构上仍然支离破碎,异构工具产生的输出难以比较、追踪和复用。基于AIRS活动的程序性条件以及《人工智能法案》对高风险系统的义务,我们为在AIRS内实现技术测试操作化的基础设施推导出11项架构和治理要求。针对这些要求,我们提出了AI评估沙箱配置器,这是一个开源框架,结合了通过稳定插件API访问的精选测试和控制目录、协调异构输出的共享数据模型、用于多学科解释的特定角色仪表板,以及面向受众的分段报告。我们描述了该框架的架构和当前版本,并报告了一项早期试点,该试点在真实的AIRS活动中演练了协调和报告层,并为官方退出报告做出了贡献。我们讨论了路线图、目录分层贡献模型引发的治理问题,以及开源评估生态系统可能在各成员国间形成的制度路径。
英文摘要
The EU's Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured technical testing, running such testing at scale demands dedicated infrastructure, yet the tooling ecosystem remains structurally fragmented, with heterogeneous tools producing outputs that are difficult to compare, trace, and reuse. From the procedural conditions of AIRS engagements and the AI Act obligations for high-risk systems, we derive 11 architectural and governance requirements for the infrastructure that operationalises technical testing within an AIRS. In response to these requirements, we introduce the AI Assessment Sandbox Configurator, an open-source framework combining a curated Catalogue of tests and controls accessed through a stable plug-in API, a shared data model that harmonises heterogeneous outputs, role-specific dashboards for multi-disciplinary interpretation, and audience-segmented reporting. We describe the architecture and current release, and report an early-stage pilot that exercised the harmonisation and reporting layers within a live AIRS engagement and contributed to an official Exit Report. We discuss the roadmap, the governance questions raised by the Catalogue's tiered contribution model, and the institutional pathways through which an open-source assessment ecosystem could emerge across Member States.