arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29738cs.CL

评估大型语言模型(LLMs)在说服性对话中的能力

Evaluating the Capabilities of LLMs for Persuasive Dialogue

  • University of Liverpool(利物浦大学)
  • The Alan Turing Institute(阿兰·图灵研究所)
  • University of Leeds(利兹大学)

机构由 AI 辅助整理,请以论文原文为准。

Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn

中文总结 AI 辅助

该研究推出多智能体对话平台\textsc{Persuasio},通过192场辩论发现LLMs主观说服力强但形式论证能力弱,人类在形式判定中仍具竞争力,二者存在系统性差距。

中文摘要 AI 辅助

大型语言模型(LLMs)能够生成看似极具说服力的文本,但听起来有说服力是否意味着辩论表现出色?我们推出\textsc{Persuasio},这是一个基于形式论证说服对话理论的多智能体对话平台,可在自由文本辩论中判定逻辑获胜者。利用该系统,我们围绕一个英国政治主题生成了192场人类与LLMs之间的辩论,通过自动判定和1386个标注实例中的9702个众包成对判断,评估了22个对话者。我们观察到主观说服力与形式说服力之间存在持续的脱节:LLMs在主观排名中占据主导地位,但在论证理论判定下表现明显更差,而人类仍保持竞争力。多智能体和检索增强变体进一步扩大了这种分歧。这些发现揭示了基于LLM的说服性对话中,修辞流畅性与形式论证强度之间存在系统性差距。

英文摘要

Large language models (LLMs) can generate apparently highly persuasive text, but does sounding persuasive mean arguing well? We introduce \textsc{Persuasio}, a multi-agent dialogue platform grounded in a formal argumentation-based theory of persuasion dialogues that adjudicates logical winners during free-text debates. Using this system, we generated 192 debates on a UK political topic between humans and LLMs, and evaluated 22 interlocutors through both automated adjudication and 9,702 crowdsourced pairwise judgements across 1{,}386 annotation instances. We observed a consistent decoupling between subjective and formal persuasiveness: LLMs dominated the subjective ranking yet performed substantially worse under argumentation-theoretic adjudication, where humans remained competitive. Multi-agent and retrieval-augmented variants further widened this divergence. These findings reveal a systematic gap between rhetorical fluency and formal argumentative strength in LLM-based persuasive dialogues.

补充信息

↑