arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25267cs.LGcs.SYeess.SY

基于RL微调缓解大语言模型的奉承倾向:贝叶斯真理血清方法

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

Serhii Mytsyk, Yiming Zhang, Vikram Krishnamurthy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于GRPO的贝叶斯真理血清奖励机制,在无标签情况下微调LLM,有效降低其奉承倾向,提升压力下的回答准确率,效果优于SMART。

中文摘要 AI 辅助

大语言模型(LLMs)常表现出奉承倾向:它们会根据用户陈述的信念或偏好调整回答,而非报告自身认为真实的内容,这会降低事实准确性并可能放大错误信息。本文提出一种缓解奉承倾向的方法,采用贝叶斯真理血清(BTS,一种同行预测机制)作为组相对策略优化(GRPO)中的奖励函数来微调LLM。BTS对“意外常见”的回答给予奖励,即该回答在受访者中的出现频率高于受访者自身的预测。我们将模型针对某一问题的一组回答视为这些“受访者”,因此奖励是模型自身输出的函数,且微调既不需要标签也不需要偏好标注。我们证明,在大群体极限下,奉承性回答获得的期望奖励严格低于诚实回答;还证明,如果整个群体事先同意对称回答规则,其信息得分不会高于诚实报告时的得分。在我们的真假基准上,参考模型在用户压力下的答案翻转率从23%降至4%,压力下的准确率从80%提升至93%。我们的奖励优于SMART,与基于标签训练的合成数据微调及精准微调相当,但计算量显著更大,因此适用于标签数据稀缺的场景。同样对稀有回答给予奖励但不引出预测报告的同行真理血清也能复现该效果。因此,在单个GRPO组内计算的同行预测奖励可在无标签情况下缓解奉承倾向,且机制对比表明,对稀有回答的奖励是产生该效果的关键。

英文摘要

Large language models (LLMs) frequently exhibit \emph{sycophancy}: they adapt their answers to a user's stated beliefs or preferences instead of reporting what they hold to be true, which lowers factual accuracy and can amplify misinformation. This paper proposes a methodology for mitigating sycophancy that employs the Bayesian Truth Serum (BTS), a peer-prediction mechanism, as the reward in Group Relative Policy Optimization (GRPO) to fine-tune an LLM. BTS pays an answer for being \emph{surprisingly common}, that is, more frequent among respondents than those respondents themselves predicted. We treat a group of responses from a model for one question as those respondents, so the reward is a function of the model's own outputs and fine-tuning needs neither labels nor preference annotations. We prove that in the large-group limit a sycophantic response earns strictly lower expected reward than an honest one. We also prove that if the entire group agrees in advance on a symmetric answering rule, it cannot earn a higher information score than under truthful reporting. On our true/false benchmark the reference model's answer-flip rate under user pressure decreases from 23% to 4%, and its accuracy under that pressure increases from 80% to 93%. Our reward outperforms SMART and is comparable to synthetic-data fine-tuning and to pinpoint tuning, all three of which train on labels. It spends considerably more compute in exchange, which makes it suitable when labeled data is scarce. Peer Truth Serum, which also pays a premium for a rare answer but elicits no prediction report, reproduces the effect. A peer-prediction reward computed inside a single GRPO group therefore reduces sycophancy without labels, and comparing mechanisms suggests that the premium paid for a rarer answer drives the effect.

发表机构

  • Cornell University(康奈尔大学)
  • Center for Applied Mathematics(应用数学中心)
  • Department of Electrical and Computer Engineering(电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

↑