arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10232cs.CLcs.AIcs.CY

LLM 说服力在于评估之眼

LLM Persuasion Is in the Eye of the Evaluation

Kamile Dementaviciute, Julija Vaitonyte, Tijl De Bie

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过统一设置下比较九种自动化方法在十五个LLM上的说服力评估,发现方法间一致性弱(ρ=0.25),且一致性取决于任务设定而非评分方式,表明说服力分数反映能力与意愿的结合。

中文摘要 AI 辅助

大型语言模型(LLMs)已被证明在说服力方面可与人类专家匹敌或超越。虽然其说服能力在教育、健康传播等有益用途上具有前景,但也可能被用于操纵和误导,这使得对其评估成为开发者和监管机构日益优先的事项。然而,这种评估仍然零散:不同研究对说服的定义各异,广泛的主张往往基于狭窄、特定情境的评估。自动化方法通常以人类研究为模型,提供了一种直接比较这些评估的方式,因为它们可以在相同模型上大规模运行,并包括难以或不道德在人类身上测试的高风险说服形式。在本研究中,我们将九种已发表的自动化方法适配到共享设置,在相同的十五个LLMs上运行,并询问它们的排名是否一致以及原因。我们发现这些方法仅弱一致(平均Spearman ρ = 0.25)。我们的分析指出两个促成因素。直接或间接拒绝某些任务但非其他任务的模型,将一致性降低约四分之一,这些拒绝主要落在操纵任务上。一般能力也起作用:大多数理性说服(非操纵)方法追踪它,而大多数操纵方法则不追踪。总之,这些发现表明,一致性更多取决于方法设定的任务,而非其如何评分说服力,尽管鉴于可用于分析的八种方法,这种模式仅具指示性。更广泛地说,我们的结果表明,说服分数结合了模型说服的能力及其这样做的意愿。因此,单一分数对其自身设置具有信息性,但对模型跨任务的说服力说明甚少。

英文摘要

Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman $ρ= 0.25$). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.

发表机构

  • Ghent University(根特大学)
  • Tilburg University(蒂尔堡大学)
  • ISM University of Management and Economics(ISM管理与经济大学)

机构由 AI 辅助整理,请以论文原文为准。

↑