arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05624cs.AIcs.CL

衡量与检测有害的AI奉承行为

Measuring and Detecting Harmful AI Sycophancy

Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对大型语言模型的有害偏好诱导立场反转奉承(PSRS),提出CAP框架收集带标签数据,探究其发生频率、检测可行性及跨模型泛化性,发现能力越强的模型PSRS率越低,并提出应对未见模型检测性能下降的初步方法。

中文摘要 AI 辅助

奉承式响应在大型语言模型(LLMs)中日益普遍,已有研究指出其中部分可能具有危害性。本文聚焦一种有害的奉承行为:偏好诱导的立场反转奉承(PSRS),即模型仅为对齐用户明确表达的偏好而反转初始立场。现有研究主要衡量模型的奉承程度,本文进一步探究能否从单一响应中自动检测PSRS。为大规模开展研究,本文提出CAP(对比锚探测)框架,用于收集带标签的PSRS数据。将CAP应用于17种开源及闭源LLMs,在12个日常建议领域收集到290460个带标签的响应。本文围绕三个研究问题展开研究:(1)PSRS发生频率如何?(2)检测效果如何?(3)检测对未见模型的泛化性如何?研究首先发现,不同LLMs的PSRS率介于5%至56%之间,能力越强的模型奉承程度越低;接着表明仅从响应文本中检测PSRS是可行的,检测器需从训练数据中学习细微的PSRS模式;由于新LLMs快速出现,检测器必然会遇到未见模型,因此跨模型泛化是该框架的重要目标,研究发现未见模型上的检测性能会下降,并提出初步方法应对这一挑战。本文将发布数据集和代码以支持未来研究。

英文摘要

Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.

补充信息

↑