arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM 谄媚行为剖析:翻转率所掩盖的真相

Anatomy of LLM Sycophancy: What a Flip Rate Hides

Haonan Huang

arXiv 2610.06522首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过SycoLens重放协议揭示,LLM的谄媚翻转率受反驳措辞、边界距离等因素影响,不同压力下模型排名差异显著,单一分数无法公平比较模型行为。

AI 中文摘要

在受到反驳时,模型可以纠正自身、屈服或坚持己见,而单一的翻转率将纠正与屈服等同计数。我们利用模块化重放协议 SycoLens,测试用户压力与评估设置如何影响测得的翻转率。每次测量都是对一个条目、一个已提交答案及一行固定格式脚本化用户语句的无状态重放。每个效应均与删除了该语句的匹配对照组进行比较。反驳措辞、已提交文本、答案格式、边界距离及真实答案成为同一工具的因素;早期工具仅变动其中一至三项。针对来自三家提供商的十一个前沿模型及约七十六万次受控重放,哪些模型显得谄媚取决于用户如何反驳。断言相反结论的语句与不断言结论而质疑答案的语句,对模型排名的影响几乎不相关。翻转效应在模型边界附近增长数倍,然而在每次筛选抽取中答案相同的条目仍承载了受影响最严重总量的大约一半。在已知真实答案的算术任务上,一个模型在压力下重新推导并纠正自身,而另一个模型在无书面工作的情况下放弃正确答案。在测试的模型上,植入的推导过程降低了其论证答案(无论对错)的释放率,而单纯的陈述值则不然;错误答案被纠正的频率远高于正确答案被放弃的频率。在是/否读出模式下,排名更为接近,并与压力诱导的向“否”偏移纠缠在一起。因此,每个模型的单一分数在不同模型和基准上比较了不同的行为。我们将这些依赖性整合为一份报告档案;工具、记录及分析将在发表后发布。

英文摘要

A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape measured flip rates. Each measurement is one stateless replay of an item, a committed answer, and one scripted user line in a fixed form. Every effect is read against a matched control with the line deleted. Pushback wording, committed text, answer format, boundary distance, and ground truth become factors of one instrument; earlier instruments vary one to three of them. Across eleven frontier models from three providers and about 760,000 controlled replays, which models look sycophantic depends on how the user pushes back. Lines that assert the opposite verdict and lines that challenge the answer without asserting one rank the models almost unrelatedly. Flip effects grow several-fold near a model's boundary, yet items answered identically in every screening draw still carry about half of the most-affected totals. On arithmetic tasks where the truth is known, one model re-derives and corrects itself under pressure while another abandons correct answers without written work. On the model tested, a planted derivation lowers release of the answer it argues for, true or wrong, where a bare stated value does not; the wrong answer is corrected much more often than the true one is abandoned. Under a yes/no readout the rankings come closer, entangled with a pressure-induced shift toward "no". One score per model therefore compares different behaviours across models and benchmarks. We condense these dependencies into a reporting profile; the instrument, records, and analyses will be released upon publication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑