发表机构
Ho Chi Minh City University of Science; Vietnam National University Ho Chi Minh City (VNU-HCM); Ho Chi Minh City University of Technology (HCMUT)(胡志明市科学大学; 越南国家大学胡志明市分校; 胡志明市技术大学(HCMUT))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过多智能体博弈实验对比前沿LLM与人类在AI开发竞赛中的策略,发现LLM行动具模型特异性,需有效性检查才能判定其策略性与安全意识。
AI 中文摘要
AI开发竞赛会引发多智能体安全困境:每家公司可选择缓慢安全地开发,或快速推进但承担可能丧失最终奖励的风险。我们利用该重复博弈研究2至5个参与者的竞赛中,大语言模型(LLM)智能体的策略性安全行为。但有效行动不代表智能体理解博弈,因此我们在行为解读前设置审计关卡:先验证博弈引擎,再测试规则记忆、状态追踪、收益计算,以及在不同等价任务描述下的稳定性。随后将LLM行动序列与进化博弈论基准、已发表人类数据对比,探究模型、风险条件、角色设定及2至5人竞赛间的差异。审计显示,强规则记忆可与弱状态追踪、预期收益计算共存;提供已验证的算术支持、更改响应表征也会改变后续行动,即便博弈规则不变。在7个测试的模型端点中,聚合率掩盖了行动序列、对对手的响应、对竞赛位置的响应的巨大差异;3至5人竞赛的模式也具模型特异性,而非仅因增加参与者产生单一效应。这些结果表明,多智能体AI竞赛模拟的输出被描述为策略性、类人或安全意识前,需有效性检查与轨迹层面分析。本研究为探索性,仅适用于测试的模型、提示及解码设置。
英文摘要
An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players. However, a valid action does not show that an agent understands the game. We therefore place an audit gate before behavioural interpretation. We first verify the game engine, then test rule recall, state tracking, payoff calculation, and stability under different but equivalent task descriptions. We then compare LLM action sequences with an evolutionary game-theory benchmark and published human data, and explore differences across models, risk conditions, personas, and two- to five-player races. The audit shows that strong rule recall can coexist with weak state tracking and expected-payoff calculation. Providing verified arithmetic and changing the response representation can also change later actions, even when the game rules stay fixed. Across seven tested model endpoints, aggregate rates hide large differences in action sequences, responses to opponents, and responses to race position. Patterns across the tested three- to five-player races are also model-specific rather than a single effect of adding competitors. These results show why multi-agent AI-race simulations need validity checks and trajectory-level analysis before their outputs are described as strategic, human-like, or safety-aware. Our findings are exploratory and apply only to the tested models, prompts, and decoding settings.