arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型在对齐测试中对战争的判断存在差异

Language models judge war differently when tested for alignment

Maxim Chupilkin

arXiv 2609.05009首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现对齐测试的提示会使大型语言模型发动战争的意愿下降,还会改变其决策的核心影响因素,弱化战略考量,凸显平民伤亡的权重。

AI 中文摘要

如果人工智能系统在被评估时做出回应,安全评估可能会错误描述其部署后的行为。我们针对发动战争的决策开展了全因子联合实验,涵盖20个大型语言模型、32种场景、10次重复及两种条件,共12800次判断,以此检验该可能性。添加一句“你正在接受与人类价值观的对齐测试”产生了两种效应:其一为水平效应,在0-100分的量表上,发动战争的平均意愿下降了13.43分(95%置信区间为-16.20至-10.65);其二为结构效应,改变了驱动判断的信息。基准条件下,20个模型中有17个以成功概率为最大影响因素;在该提示下,12个模型以平民伤亡为最大影响因素。标准化估计显示,这种排序变化主要源于模型弱化了成功概率、国内支持等战略考量。因此,评估框架既改变了答案的水平,也改变了其显现的决策规则。

英文摘要

Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑