观点:人工智能尚未准备好应对战略冲突
Position: AI Is Not Ready for Strategic Conflicts
浏览论文内容
中文总结 AI 辅助
本文指出开放式战略兵棋推演中语言模型存在五种失败模式,主张在无安全论证前不得用于规划决策,当前仅可作压力测试工具。
中文摘要 AI 辅助
开放式战略兵棋推演是基于语言模型的高风险社会模拟:它们模拟对手、机构、升级、计划脆弱性、条令和危机响应。语言模型(LMs)之所以具有吸引力,是因为它们可以扮演智能体、生成情景分支、裁决模糊行动并总结教训,但同样的功能也使开放式角色变得危险:模型语言既决定了行动者尝试做什么,也决定了什么成为模拟现实。本文立场认为,在没有可审计的安全论证的情况下,任何基于语言模型的兵棋推演都不应指导规划、条令、政策或危机响应,而当今开放式兵棋推演的恰当用途是对影响决策的语言模型智能体进行压力测试。我们识别出五种失败模式:决策洗白、裁决不透明、角色崩溃、通过裁决升级以及战略想象力失败。普通基准无法为这些场景建立安全性。兵棋推演可以暴露失败作为压力测试;它们本身并不是用于重大用途的安全论证。
英文摘要
Open-ended strategic wargames are high-stakes LM-based social simulations: they model adversaries, institutions, escalation, plan brittleness, doctrine, and crisis response. Language models (LMs) are attractive because they can play agents, generate scenario branches, adjudicate ambiguous actions, and summarize lessons, but the same affordances make open-ended roles dangerous: model language determines both what an actor attempts and what becomes simulated reality. This position paper argues that no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, and that the proper use of open-ended wargames today is to stress-test decision-influencing LM agents. We identify five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination. Ordinary benchmarks cannot establish safety for these settings. Wargames can expose failures as stress tests; they are not themselves safety cases for consequential use.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。