arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁会受到扣留延迟的影响?非对称扩散下开放权重人工智能发布的博弈论模型

Who Does Withholding Delay? A Welfare Model of Open-Weight AI Release Under Asymmetric Proliferation

Daniel Commey

arXiv 2607.22957首次发表:更新:

AI 中文总结

研究非对称扩散下开放权重人工智能发布,通过博弈论建模,探讨不同发布方式及影响因素,如访问反转和非对称赋权,给出政策排名相关因素,确定发布审查应估计的数量,为相关决策提供依据。

AI 中文摘要

限制对两用人工智能模型的访问只有在对有害行为者的延迟超过防御者时才具有预防性。这种情况因行为者而异:国家机构或有组织犯罪集团可能通过盗窃、提炼、中介访问、自主开发或国外发布获得替代品,而小型公用事业公司或开源维护者可能没有可比途径。我们对实验室在受控访问、防御者优先窗口、受保护的开放权重和最低限制的开放权重之间的选择进行建模。当限制赋予对手比防御者更快获得有效替代品的访问优势时,就会出现访问反转。当立即发布为最不可能拥有替代品的群体增加最多能力时,就会出现非对称赋权。政策排名还取决于相对有用性、机会主义滥用、攻防转换、防御溢出、保障摩擦和不可召回损失。线性基准产生一个独特的对手替代阈值,当端点条件成立时,超过该阈值广泛发布将超过控制。当选定的防御者在对手赶上之前部署保护时,防御者优先窗口具有价值,当可移除的保障措施能够阻止足够的机会主义滥用时,它们仍然有用。非线性实现为每个发布层提供一个非空的政策区域。三个嵌套的2048点确定性设计评估对参数边界的敏感性,一个单独的网格检查发布后特定行为者的部署延迟。发布、网络评估和事件响应案例确定发布审查应估计的数量:特定行为者的替代时间、边际能力增益、部署率、防御范围、新出现的滥用和不可召回损失。

英文摘要

Withholding a dual-use AI model delays only the actors that lack other routes to a comparable capability. If sophisticated adversaries obtain substitutes faster than distributed defenders, restriction can delay defenders more than the adversaries it targets. We compare controlled access, a defender-first window followed by public release, safeguarded open weights, and minimally restricted open weights in a discounted welfare model with actor-specific substitute acquisition. Under exponential acquisition, restriction gives adversaries a positive discounted access advantage exactly when they substitute faster than defenders, and, with equal usefulness, immediate release adds more expected capability at a fixed horizon to the slower-substituting group. Neither result implies that release is preferable, because opportunistic misuse, defensive reach, safeguard friction, and irreversible losses can reverse the ranking. In a linear benchmark, broad release overtakes control above a unique adversary-substitution threshold whenever such a threshold exists, and we derive the probability that selected defenders deploy before both adversary substitution and public release. In a nonlinear implementation, each of the four policies is optimal somewhere in the parameter space. Three nested 2,048-point designs over thirteen inputs show that policy shares depend strongly on the chosen parameter bounds. Release records and cybersecurity reports illustrate the quantities a release review would need to measure and are kept separate from the calibration.

Comments22 pages, 8 figures, 5 tables. Code and data are available at https://github.com/dcommey/asymmetric-proliferation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑