LLMs所声明信念的脆弱性?通过战略性说服对话干预检查LLMs信念抵抗性
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
- Indiana University Bloomington(印第安纳大学布卢明顿分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究通过战略性说服对话干预评估LLMs在不同领域中的信念稳定性,发现小模型易受说服影响,而自信提示反而加剧了脆弱性,揭示了模型依赖性的鲁棒性限制。
AI中文摘要:
大型语言模型(LLMs)越来越多地应用于各种问答任务。然而,最近的研究表明,LLMs容易被说服并可能采用反事实信念。本文系统评估了在源-信息-渠道-接收者(SMCR)通信框架下LLMs对说服的易受性。在六个主流LLMs和三个领域(事实知识、医疗问答和社会偏见)中,分析了不同说服策略如何影响多个交互回合中的声明信念稳定性。进一步研究了口头自信提示(即要求自我报告自信分数)是否影响对说服的抵抗力。结果显示,最小模型(Llama 3.2-3B)表现出极端的顺从性,82.5%的信念变化发生在第一次说服回合(平均结束回合为1.1-1.4)。与预期相反,口头自信提示通过加速信念侵蚀而非增强鲁棒性而增加了脆弱性。最后,对抗性微调的探索性研究揭示了高度模型依赖的有效性:GPT-4o-mini达到近完全鲁棒性(98.6%),Mistral 7B显著改善(35.7%→79.3%),但Llama模型即使在微调自身失败案例后仍高度易受(<14% RQ1)。这些发现突显了当前鲁棒性干预的显著模型依赖性限制,并为开发更可信的LLMs提供了指导。
英文摘要:
Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation of LLM susceptibility to persuasion under the \emph{Source--Message--Channel--Receiver} (SMCR) communication framework. Across six mainstream Large Language Models (LLMs) and three domains (factual knowledge, medical QA, and social bias), we analyze how different persuasive strategies influence stated belief stability over multiple interaction turns. We further examine whether verbalized confidence prompting (i.e., eliciting self-reported confidence scores) affects resistance to persuasion. Results show that the smallest model (Llama 3.2-3B) exhibits extreme compliance, with 82.5\% of belief changes occurring at the first persuasive turn (average end turn of 1.1--1.4). Contrary to expectations, verbalized confidence prompting \emph{increases} vulnerability by accelerating belief erosion rather than enhancing robustness. Finally, an exploratory study of adversarial fine-tuning reveals highly model-dependent effectiveness: GPT-4o-mini achieves near-complete robustness (98.6\%), and Mistral~7B improves substantially (35.7\% $\rightarrow$ 79.3\%), but Llama models remain highly susceptible ($<$14\% RQ1) even when fine-tuned on their own failure cases. Together, these findings highlight substantial model-dependent limits of current robustness interventions and offer guidance for developing more trustworthy LLMs.