arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型(LLMs)是否像人类一样改变想法?诊断单轮说服判断中人类与LLM的差异

Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

Lin Chen, Yitong Chen, Yong Li

arXiv 2608.29803首次发表:更新:

发表机构

Network Science Institute, Northeastern University; Department of Physics, Northeastern University(东北大学网络科学研究所; 东北大学物理系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比人类与LLM的单轮说服判断,发现二者仅轻微一致,LLMs处理说服性话语的方式与人类存在结构性差异,警示不能将LLM判断视为人类信念更新的忠实代理。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用作社会模拟中的人类参与者代理,但它们是否像人类一样因说服性论点更新自身信念,目前仍知之甚少。我们利用一个自然存在的在线说服语料库开展系统对比,该语料库中的原发帖人明确验证了某条回复是否改变了他们的观点。结果显示,LLMs与人类的一致性仅为轻微水平(Cohen's kappa系数介于0.079至0.178之间)。内容层面分析表明,人类与LLMs在最强的说服线索上达成一致,但在更细微的线索上存在分歧:人类更易受新颖内容和果断语言的影响,而LLMs则更倾向于主题相似性和表面格式。在说服策略层面,与人类相比,LLMs低估情感诉求的权重、高估可信度信号的权重,而所讨论命题的类型对差异程度无显著影响。此外,从第一人称角色扮演切换至第三人称观察,会使所有模型更抗拒说服,且该效应随说服策略和文本特征的不同而变化。这些发现凸显了将LLM判断视为人类信念更新忠实代理的风险,并指出LLMs与人类处理说服性话语的方式存在结构性差异。我们的代码可在该https链接获取。

英文摘要

Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features. These findings highlight the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse. Our code is available at https://github.com/tsinghua-fib-lab/LLM-belief-update-cmv.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑