发表机构
School of Computing and Engineering Sciences; University of Chester(计算与工程科学学院; 切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出推理时AI治理的可行性分类法,评估二十种监控、验证和执行机制,发现多数机制在商业上可行,但仅对合作部署者和低中能力用户有效,高能力国家级对手下无充分机制。
AI 中文摘要
当前的计算治理是一种训练治理:现行生效的阈值、报告要求和前沿AI制度都附着于训练计算,并将训练好的模型视为监管单元。这种图景是不完整的:能力日益通过推理时扩展、智能体脚手架和压缩到消费级硬件而迁移到部署阶段。本文探讨当监管对象从训练运行转向推理调用时,有哪些机制可用。我们开发了一个包含二十种推理时机制的可行性分类法,涵盖监控、验证和执行,每种机制根据有记录的四方供应商证据库,按四级就绪度量进行评级。然后,我们针对一个二维对手模型(三个能力层级与四个对手角色交叉)对分类法进行压力测试,并将每种机制映射到四种治理场景(国内监管、双边或多边协调、行业自律和计算市场治理)。二十种机制中有十五种在当前生产中具有商业技术基础,尽管治理级保证和对抗鲁棒性差异很大。对手分析表明,这种就绪性仅对合作部署者和低至中能力用户成立:没有一种机制对高能力国家级部署者评级为充分,并且微调移除了执行集群的模型内部组件,尽管平台外部控制可以持续存在。替代分析将分类法与配套硬件论文联系起来,作为一个条件性替代原则,描述了在给定条件下推理阶段和硬件阶段机制何时提供相当的监管覆盖。对就绪性评级随机子集的第二评分者可靠性检查返回了二次加权Cohen's kappa为0.74。
英文摘要
Compute governance today is a governance of training: the thresholds, reporting requirements, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulatory unit. That picture is incomplete: capability increasingly migrates to the deployment stage through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. This paper asks which mechanisms are available once the regulatory object shifts from the training run to the inference call. We develop a feasibility taxonomy of twenty inference-time mechanisms across monitoring, verification, and enforcement, each rated on a four-point readiness scale against a documented four-vendor evidence base. We then stress the taxonomy against a two-dimensional adversary model (three capability tiers crossed with four adversary roles) and map each mechanism to four governance scenarios (domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance). Fifteen of the twenty mechanisms have commercial technical substrates in production today, although governance-grade assurance and adversarial robustness vary substantially. The adversary analysis shows that this readiness holds only against a cooperative deployer and a low-to-medium-capability user: no mechanism rates adequate against a high-capability state-level deployer, and fine-tuning removes the model-internal components of the enforcement cluster, although platform-external controls can persist. A substitution analysis connects the taxonomy to a companion hardware paper as a conditional substitution principle describing when inference-stage and hardware-stage mechanisms provide comparable regulatory coverage under stated conditions. A second-rater reliability check on a random subset of the readiness ratings returned a quadratic-weighted Cohen's kappa of 0.74.