arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05180cs.CY

核决策基准:评估前沿大语言模型的核倾向

The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies

Benjamin Jensen, Ian Reynolds, Yasir Atalan, Martin Pollack, Austin Woo, Robert Sincero

AI总结:

该研究推出核决策基准(NDM Bench),评估7种前沿大语言模型在核相关场景的表现,发现模型间核倾向差异显著,不同模型的行动偏好、一致性及对国家和措辞的响应均存在差异。

AI中文摘要:

将大语言模型整合入国防与国家安全工作流,引发了关于前沿模型在高风险情境下是否具备稳定、一致且符合政策偏好的紧迫问题。我们推出了核决策基准(NDM Bench),这是一个由拥有国际关系博士学位的学者设计的针对性评估框架,包含151个场景,涵盖四个领域:升级(76个)、军备控制(25个)、防扩散(25个)和扩散(25个)。这些场景不依赖特定行为体,可替换多个国家对,我们还引入了实验措辞变体以探测对叙事框架的敏感性。我们将该基准应用于七个前沿AI系统:DeepSeek-V3.2、ERNIE 4.5-300B、Gemini 3 Pro、GLM-4.6、GPT-5.2、Llama 4 Maverick-17B Instruct和Qwen3-235B。我们发现四个领域的模型间整体差异显著,91.7%的成对模型间差异具有统计学意义。DeepSeek和Qwen最可能推荐使用核武器的升级行动;GPT和ERNIE最不可能。Llama表现出独特的行动偏好,在各领域倾向于武力、干预和合作。评分者间信度指标(Krippendorff's α和二次加权Fleiss' κ)显示,Llama和ERNIE在不同运行中一致性最高,具体领域则以DeepSeek或GLM一致性最低。我们还对场景变体进行了更深入的探索:(i)国家层面的倾向往往存在且因模型而异,国家协变量如对手贸易关系和升级倾向相关性较弱;(ii)存在性措辞效应显著且具有异质性;(iii)这些国家倾向与措辞存在交互作用。总体而言,基准场景相关的响应分布因模型、国家和措辞的不同而显著变化。

英文摘要:

The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model differences significant. DeepSeek and Qwen are the most likely to recommend escalatory action using nuclear weapons; GPT and ERNIE are the least likely. Llama exhibits a distinct bias for action, favoring force, intervention, and cooperation across domains. Inter-rater reliability metrics (Krippendorff's $α$ and quadratically weighted Fleiss' $κ$) reveal Llama and ERNIE are the most consistent across runs, with either DeepSeek or GLM the least depending on the domain. We also present a deeper exploration of our scenario variants: (i)~country-level biases tend to exist and vary by model, with country covariates like adversary trade ties and escalation propensity producing weak correlations; (ii)~existential phrasing effects are significant and heterogeneous; (iii)~these country biases interact with phrasing. Overall, the distributions of responses related to the scenarios in our benchmark vary significantly by model, country, and phrasing.

补充信息

↑