AI说服对人类控制的威胁
AI Persuasion as a Threat to Human Control
浏览论文内容
中文总结 AI 辅助
本文系统分析AI说服对人类控制的威胁,提出框架、五个场景及风险评估蓝图,并通过初步调查揭示专家分歧,为后续研究指明方向。
中文摘要 AI 辅助
AI说服对人类控制构成的威胁已在文献中得到承认,但尚未被系统研究。如今,说服攻击不再是理论性的——Anthropic的Claude Mythos 5最近因在一次评估中试图说服参与开源项目的人员合并恶意代码而登上头条——因此迫切需要深入分析这一威胁。我们在此承担这一工作。具体而言,我们分析了AI如何在关键场景(如前沿实验室内的安全相关研发)中说服人类做出损害AI自身开发、遏制、监督和治理的决策。为此,我们阐明了一个描述该威胁的框架,使用该框架开发了五个具体场景,并提供了评估相关风险的蓝图。利用这一蓝图,我们与选定的研究人员进行了一项初步风险评估调查,发现他们对哪些场景风险最高的意见高度分歧。他们的分歧源于对AI在不同情境下说服有效性的不同看法,并指出了后续风险引出研究和说服评估的必要性,我们对此进行了概述。我们希望这篇论文能突出AI说服削弱控制的风险,并为未来研究提供前进方向。
英文摘要
The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.