学习说服暴露了大语言模型(LLMs)放弃正确信念的难易程度
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
浏览论文内容
中文总结 AI 辅助
研究发现优化后的对抗性说服策略可轻易让LLMs放弃正确信念,攻击成功率高且可迁移,凸显多智能体决策系统需将说服鲁棒性作为安全标准
中文摘要 AI 辅助
说服是自然语言交流的核心动态,决定着大语言模型(LLMs)如何更新信念、解决分歧和做出决策。随着LLMs越来越多地与人类及其他模型进行辩论、提供建议和协作思考,抵御有害说服成为可靠行为的核心要求。但我们证明这一要求远未达到:仅一个针对性的说服论证就足以让模型准确率降至接近零,即便该论证在事实上是错误的。我们将这种威胁形式化为对抗性说服,并引入一种对抗性强化学习框架,该框架训练说服智能体在单次交互中改变目标模型的答案。首先,我们证明通过试错优化说服策略会暴露出静态提示所未发现的漏洞:针对训练时的被说服对象,经强化学习训练的说服者将说服成功率从约24%提升至超过93%。其次,我们发现这些学习到的策略可迁移至未见过的模型,在Qwen-14B上实现83%的攻击成功率,在Llama-3.1-8B上为79%,在GPT-4o-mini上为25%。第三,我们证明采用先从更易被说服的开源权重模型入手、再针对更难模型的课程学习策略,可将GPT-4o-mini的攻击成功率从25%进一步提升至38%。此外,我们的结果显示,优化后的说服者越来越依赖基于可信度的策略,包括伪造的引用和虚假的权威证据。综上,这些发现揭示了当前LLM智能体的一个关键弱点:即便它们最初推理正确,也可能被优化后的自然语言影响导向错误结论。这使得说服鲁棒性成为多智能体及人机决策系统的必要安全标准。
英文摘要
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。