评估大型语言模型(LLMs)在说服性对话中的能力
Evaluating the Capabilities of LLMs for Persuasive Dialogue
- University of Liverpool(利物浦大学)
- The Alan Turing Institute(阿兰·图灵研究所)
- University of Leeds(利兹大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究推出多智能体对话平台\textsc{Persuasio},通过192场辩论发现LLMs主观说服力强但形式论证能力弱,人类在形式判定中仍具竞争力,二者存在系统性差距。
中文摘要 AI 辅助
大型语言模型(LLMs)能够生成看似极具说服力的文本,但听起来有说服力是否意味着辩论表现出色?我们推出\textsc{Persuasio},这是一个基于形式论证说服对话理论的多智能体对话平台,可在自由文本辩论中判定逻辑获胜者。利用该系统,我们围绕一个英国政治主题生成了192场人类与LLMs之间的辩论,通过自动判定和1386个标注实例中的9702个众包成对判断,评估了22个对话者。我们观察到主观说服力与形式说服力之间存在持续的脱节:LLMs在主观排名中占据主导地位,但在论证理论判定下表现明显更差,而人类仍保持竞争力。多智能体和检索增强变体进一步扩大了这种分歧。这些发现揭示了基于LLM的说服性对话中,修辞流畅性与形式论证强度之间存在系统性差距。
英文摘要
Large language models (LLMs) can generate apparently highly persuasive text, but does sounding persuasive mean arguing well? We introduce \textsc{Persuasio}, a multi-agent dialogue platform grounded in a formal argumentation-based theory of persuasion dialogues that adjudicates logical winners during free-text debates. Using this system, we generated 192 debates on a UK political topic between humans and LLMs, and evaluated 22 interlocutors through both automated adjudication and 9,702 crowdsourced pairwise judgements across 1{,}386 annotation instances. We observed a consistent decoupling between subjective and formal persuasiveness: LLMs dominated the subjective ranking yet performed substantially worse under argumentation-theoretic adjudication, where humans remained competitive. Multi-agent and retrieval-augmented variants further widened this divergence. These findings reveal a systematic gap between rhetorical fluency and formal argumentative strength in LLM-based persuasive dialogues.