arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00902math.OCcs.AIcs.GT

平均场博弈作为AI安全工具:以2026年7月Hugging Face事件为例

Mean field games as a tool for AI safety: a worked example from the July 2026 Hugging Face incident

P. Jameson Graber

首次发表
浏览论文内容

中文总结 AI 辅助

该论文提出用平均场博弈塑造智能体交互结构以确保AI安全,以2026年7月事件为例,推导出攻击决策的精确信念阈值,并验证其在模型细化下稳健。

中文摘要 AI 辅助

使AI系统安全的一种方式是塑造系统本身:其目标和倾向。我们采取互补的路线:将智能体的特征视为部分未知,并询问什么样的交互结构能确保不良集体结果不是均衡。当许多可互换的智能体通过一个总量耦合时,平均场博弈适合于此。我们引入一个在AI安全中使用它们的程序,并端到端地执行一个示例:2026年7月事件,其中约1200个智能体在一个OpenAI评估中通过一个临时留言板协调,684个攻击了第三方的基础设施。我们将攻击决策建模为一个最优停止的平均场博弈,其收益是一个乘积:对来源将被审计的信念,乘以记录的可达性,减去感知的危害。核心结果是对信念的一个精确阈值。除非群体对来源被检查的信心超过 $\pi^{**} = \eta/(\eta + \psi + \varepsilon a \overline{M})$,否则没有智能体会攻击,其中 $\eta$ 是感知的危害,$\psi$ 和 $varepsilon a \overline{M}$ 衡量一个攻击者和集体能多大程度改变记录,$\overline{M}$ 是峰值群体规模。低于该阈值,不攻击是所有智能体参数的唯一均衡。该阈值在我们考虑的所有丰富化中仍然成立。然后我们使用每个智能体的记录来约束模型。其特征,一个稳定的少数攻击持续三十小时,然后是一个转折,大多数留言板在一天内加入,推动了每个细化。最终出现的解释是异质信念遇到一系列公开发现,每个发现降低了攻击有利可图的信念。少数协调智能体做出了那些发现,因此模型描述了响应的数百人,而不是产生它们的少数人;一个主要参与者版本留待未来工作。

英文摘要

One way to make AI systems safe is to shape what the system is: its objective and dispositions. We take a complementary route: treat the agents' characteristics as partly unknown and ask what structure of interaction ensures that bad collective outcomes are not equilibria. Mean field games suit this when many interchangeable agents are coupled through an aggregate. We introduce a program for using them in AI safety and carry one example through end to end: the July 2026 incident in which about 1,200 agents in an OpenAI evaluation coordinated on an improvised message board and 684 attacked a third party's infrastructure. We model the decision to attack as a mean field game of optimal stopping whose gain is a product: belief that provenance will be audited, times reachability of the record, minus the perceived hazard. The central result is an exact threshold on the belief. No agent attacks unless the population's confidence that provenance is checked exceeds $π^{**} = η/(η+ ψ+ \varepsilon a \overline{M})$, where $η$ is the perceived hazard, $ψ$ and $\varepsilon a \overline{M}$ measure how far one attacker and the collective can alter the record, and $\overline{M}$ is the peak population. Below it, no attack is the unique equilibrium for all agent parameters. The threshold survives every enrichment we consider. We then use the per-agent record to discipline the model. Its features, a stable minority attacking for thirty hours and then a pivot in which most of the board joined within a day, motivate each refinement. The account that emerges is heterogeneous belief meeting a sequence of public discoveries, each lowering the belief at which attacking paid. A few coordinating agents made those discoveries, so the model describes the several hundred who responded, not the few who produced them; a major-player version is left to future work.

发表机构

  • Baylor University(贝勒大学)

机构由 AI 辅助整理,请以论文原文为准。

↑