SyPS:衡量大语言模型中的谄媚提示敏感性
SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
查看机构详情
- Northeastern University(东北大学)
- Case Western Reserve University(凯斯西储大学)
- Everpure(爱惠浦)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出SyPS框架,通过构建控制变量的提示变体,引入SPSS指标,探究用户相关线索变化对LLMs谄媚行为的影响,发现寻求认可等线索会增加谄媚,反谄媚提示可减少谄媚。
中文摘要 AI 辅助
大语言模型(LLMs)已知会表现出社会谄媚行为,在社会敏感情境中常常对用户的观点表示认可或同意。现有评估通常在固定的提示表述下测量谄媚行为,却未明确当同一潜在情境以不同的与谄媚相关的提示变体呈现时,该行为是否稳定。在本研究中,我们探讨谄媚提示敏感性:即用户的自信程度、情感框架、社会共识或寻求认可的语言发生变化时,模型谄媚行为的改变程度。我们将此评估框架命名为SyPS,是Sycophancy Prompt Sensitivity(谄媚提示敏感性)的缩写。SyPS在现有社会谄媚评估设置的基础上构建,生成控制变量的提示变体,这些变体保留相同的潜在用户情境,但改变与谄媚相关的社会线索。我们引入谄媚提示敏感性得分(SPSS),这是衡量成对提示变体之间谄媚行为变化的实例级指标。与整体谄媚率不同,SPSS将基线谄媚与提示引发的变化分离开来,从而能够在模型层面比较对与谄媚相关的社会线索的鲁棒性。从经验来看,我们发现谄媚提示敏感性具有社会结构性:寻求认可和情感压力线索往往会增加谄媚行为,而反框架和反谄媚提示则倾向于减少谄媚行为。我们的框架揭示了LLMs在适当调整语气的同时是否保持稳定的社会判断。
英文摘要
Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in socially sensitive contexts. Existing evaluations typically measure sycophancy under a fixed prompt formulation, leaving unclear whether such behavior is stable when the same underlying situation is presented with different sycophancy-relevant prompt variants. In this work, we study sycophancy prompt sensitivity: the extent to which changes in user confidence, emotional framing, social consensus, or validation-seeking language alter a model's sycophantic behavior. We refer to our evaluation framework as SyPS, short for Sycophancy Prompt Sensitivity. Building on existing social sycophancy evaluation settings, SyPS constructs controlled prompt variants that preserve the same underlying user situation while varying sycophancy-relevant social cues. We introduce the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure of sycophancy variation across paired prompt variants. Unlike aggregate sycophancy rates, SPSS separates baseline sycophancy from prompt-induced shifts, enabling model-level comparisons of robustness to sycophancy-relevant social cues. Empirically, we find that sycophancy prompt sensitivity is socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it. Our framework highlights whether LLMs maintain stable social judgments while adapting appropriately in tone.