发表机构
Vrije Universiteit Brussel; Université Libre de Bruxelles; Teesside University(布鲁塞尔自由大学; 布鲁塞尔自由大学; 提赛德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能竞赛中速度与安全的关系,通过实验发现不安全发展受竞赛战略状态影响,引入进化模型解释,表明政策应注重降低竞争压力、促进合作,而非仅关注个体风险。
AI 中文摘要
技术竞赛在速度与安全间制造紧张关系:行动者即便冒险发展有害,但比对手更快行动可能获利。这在人工智能辩论中很突出,竞争压力常促使更冒险、安全意识更低的发展。我们通过基于理想化人工智能竞赛的框架行为实验研究此问题,配对参与者在不确定时间范围内反复在安全与不安全发展间选择。不安全发展进步更快、即时收益更高,但会累积私人风险,最高分别达10%、60%或90%,竞赛竞争结构不变,仅最大风险不同。数据不支持预先注册的风险水平比较及引出的风险偏好作用。相反,由任务重复结构激发的探索性分析表明,不安全行为受竞赛演变战略状态而非风险偏好影响更大:对手选择不安全后,参与者更可能选不安全,领先减少不安全行为,落后则增加,首轮选择可预测后续行为。为解释这些影响,我们引入简化进化模型,有始终安全、始终不安全、有条件安全和有条件反社会安全四种策略,该模型再现了处理效果并展示了竞争竞赛动态如何青睐有条件不安全行为。实验和模型共同表明,不安全发展可能源于早期行为势头、对手行为和对落后的恐惧,而非仅源于风险偏好,这表明政策应注重降低人工智能发展中的竞争压力并促进合作,而非仅关注个体风险。
英文摘要
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.
Comments45 pages (main + SI)