arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有限主权与控制成本:当部署方不拥有模型时的AI监管定价

Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model

Zhen Wen Lim

arXiv 2608.19216首次发表:更新:

AI 中文总结

本文针对部署方不拥有模型的AI监管场景,提出有限主权概念与主权折扣成本,通过135万次模拟实验明确控制协议需明确访问假设的结论。

AI 中文摘要

AI控制研究旨在解决即使模型可能存在对齐问题,如何安全部署它们的问题,但许多控制协议假设部署方可以对模型及其周边流程进行检测。对于通过API或托管端点使用前沿模型的受监管组织而言,这一假设往往不成立,此类组织中部署方可能控制业务流程,但无法掌控模型权重、服务基础设施、内部痕迹、更新流程或完整交互日志。本文引入有限主权概念,即对AI栈的数据、模型、基础设施和交互层拥有部分技术与合同层面的访问权限,认为这些访问条件决定了哪些控制协议可在实践中执行。本文贡献了四层访问类型学、按层划分的协议需求矩阵,以及主权折扣成本概念:即控制成本中用于通过合同、架构、审计、供应商保证、残余风险或缩小系统范围来替代缺失访问权限的部分。本文还报告了一项针对135万次合成案例模拟的合成访问消融实验,并通过匿名国家支付基础设施场景解读研究结果,该实验并非真实支付系统证据,而是一项构念效度研究。结果显示,完整日志可提升诊断能力,执行前网关可实现干预,痕迹访问与模型版本控制可强化事件后解释,缩小系统范围可在提升安全性的同时降低实用性。因此,被提议作为通用安全解决方案的控制协议应明确说明其访问假设。

英文摘要

AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer can instrument the model and its surrounding pipeline. That assumption often fails for regulated organisations using frontier models through APIs or managed endpoints, where the deployer may control the business process but not the model weights, serving infrastructure, internal traces, update process, or full interaction logs. This paper introduces bounded sovereignty: partial technical and contractual access across the data, model, infrastructure, and interaction layers of the AI stack. It argues that these access conditions determine which control protocols can be executed in practice. The paper contributes a four-layer access typology, a protocol-by-layer requirements matrix, and the concept of sovereignty discount cost: the part of the control tax spent substituting for missing access through contracts, architecture, audit, vendor assurance, residual risk, or reduced system scope. It also reports a synthetic access-ablation experiment over 1.35 million synthetic case simulations and interprets the findings through an anonymised national-payments-infrastructure scenario. The experiment is not real-world payment-system evidence; it is a construct-validity exercise. The results show that complete logs improve diagnosis, a pre-execution gateway enables intervention, trace access and model-version control strengthen post-incident explanation, and scope restriction can improve safety while reducing usefulness. Control protocols proposed as general safety solutions should therefore state their access assumptions explicitly.

Comments25 pages, 5 figures. Synthetic access-ablation study; not real-world payment-system evidence

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑