arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

连续训练阶段与大型语言模型说服力:错位、监督微调与偏好优化的影响

Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization

Antony Dalmiere, Pascal Marchand, Guillaume Auriol, Vincent Nicomette

arXiv 2610.09964首次发表:更新:

发表机构

CNRS; LAAS-CNRS; LERASS; Université de Toulouse(法国国家科学研究中心; 法国国家科学研究系统分析与体系结构实验室; 传播研究实验室; 图卢兹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过835名参与者的实验,发现说服性监督微调显著提升LLM说服力,而偏好优化无额外增益。

AI 中文摘要

大型语言模型(LLMs)可以被调整以影响人类态度,然而连续的后训练阶段各自的贡献仍不明确。本研究考察了三个连续训练阶段如何影响LLM的说服力:(1)通过在阴谋论数据上进行监督微调(SFT)实现错位,(2)在论证性数据上进行额外的说服性SFT,以及(3)身份偏好优化(IPO),一种偏好优化方法。在Prolific上招募的共835名参与者被随机分配到五个组间条件(中性文本、阴谋训练模型、说服训练模型、偏好优化模型和GPT-4),并在所有模型条件下,根据参与者个人档案个性化地接触了关于10个分裂性政治议题的文本。态度变化通过连续李克特量表上暴露前与暴露后立场的差异来衡量,并使用协方差分析(ANCOVA)进行分析。显著的条件×基线态度交互作用,F(4, 825)= 5.33,p < .001,表明训练效果取决于参与者的初始态度。说服性SFT产生的态度变化大于仅阴谋训练,d = 0.30,而IPO未提供额外益处,d = 0.03,且GPT-4与中性文本无差异,d = -0.01。这些结果表明,在说服性数据上进行有针对性的监督训练能提高LLM的说服力,而偏好优化在其之上未产生显著增益。

英文摘要

Large language models (LLMs) can be tuned to influence human attitudes, yet the respective contributions of successive post-training stages remain un-clear. This study examines how three successive training stages affect LLM persuasiveness: (1) misalignment through supervised fine-tuning (SFT) on conspiracy data, (2) additional persuasive SFT on argumentative data, and (3) Identity Preference Optimization (IPO), a preference-optimization method. A total of 835 participants recruited on Prolific were randomly assigned to five between-subject conditions (neutral text, conspiracy-trained model, persuasion-trained model, preference-optimized model, and GPT-4) and were exposed to texts on 10 divisive political issues, personalized from their individual profiles in all model conditions. Attitude change was measured as the difference between pre- and post-exposure positions on continuous Likert scales and analyzed with an analysis of covariance (ANCOVA). A significant condition x baseline-attitude interaction, F (4, 825) = 5.33, p < .001, indicated that training effects depended on participants' initial attitudes. Persuasive SFT produced greater attitude change than conspiracy training alone, d = 0.30, whereas IPO provided no additional benefit, d = 0.03, and GPT-4 did not differ from neutral text, d = --0.01. These results show that targeted supervised training on persuasive data increases LLM persuasiveness, whereas preference optimization yields no significant gains beyond it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑