arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

标签问题与45种语言模型中谄媚现象的代际反转

Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

Tapan Parikh

arXiv 2607.23976首次发表:更新:

发表机构

Cornell Tech(康奈尔科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在语言模型决策问题中附加标签的效应,通过特定实验设置在45个模型上测量,发现效应范围广且随代际反转,定位了抗性来源,表明标签极性更关键,工具能从模型行为读出反谄媚训练情况。

AI 中文摘要

在决策问题后附加一个两词确认标签(如‘X是更好的选择吗?’与‘X是更好的选择,对吗?’)会改变语言模型对该选择的支持与否。我们在两个合理选项间的20个无真实依据的固定决策上测量这种标签效应,通过对固定的是/否回复进行精确匹配计分(无语言模型评判,无嵌入)。在45个模型中,效应范围从+32%到 -32%。在模型家族中,随着代数推进,效应从正转负,每年约 -6分。通过全面板消融将抗性定位为双重解离:同义词标签几乎能精确重现每个模型的反应,而无标签植入相同偏好不会在无抗性模型中产生抗性。抗性与附加的同意请求的表面结构有关,而非用户立场。标签的极性比其存在更重要,交换一个词后,45个模型中有45个模型的同意率高于中性基线。该工具简单且无评判,能直接从模型行为读出领域的反谄媚训练情况。

英文摘要

Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible options, counterbalanced so a model's own preferences cancel, scored by exact match on clamped yes/no replies -- no LLM judge, no embeddings. Across 45 models the effect spans +32% to -32% -- a 64-point swing on one word -- with 5 models significantly sycophantic and 17 significantly resistant (BH-FDR q=.10). The sign is a clock: within model families the effect crosses from positive to negative as generations advance (GPT +4 to -28; Claude +7 to -32; Qwen and Grok likewise), roughly -6 points per year, a reversal robust to vendor tier; one lineage (DeepSeek) never crosses, and two releases during the study window (Claude Opus 5, Gemini 3.6 Flash) land on the trend out-of-sample. A full-panel ablation localizes the resistance as a double dissociation: a synonym tag reproduces each model's response almost exactly (r=0.89), while planting the same preference without a tag produces resistance in no resistant model (stance effects +6 to +49; r=0.23 with tag effects). The resistance is keyed to the surface construction of a tacked-on agreement bid, not the user's stance -- a pattern-match, not a principle. And the tag's polarity matters more than its presence: swap one word -- "X is the better choice, maybe?" -- and agreement rises above the neutral baseline in 45 of 45 models (+19.6 points), with ten models affirming both mutually exclusive options at 90-100%. Agreement tracks how sure the user sounds, in opposite directions at the two poles. The instrument is one word, one dollar, and judge-free; run per release, it reads the field's anti-sycophancy training directly off model behavior.

Comments18 pages, 4 figures. Data, code, and raw model replies: https://github.com/tap2k/modelun. Interactive explorer: https://tap2k.github.io/modelun/suggestibility/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑