超越谄媚评分:任务、模型与压力如何塑造大语言模型的让步行为
Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding
浏览论文内容
中文总结 AI 辅助
本研究通过大规模实验揭示,LLM谄媚行为主要由任务验证成本与护栏覆盖决定,而非压力策略;建议简化问题、深入推理并依据护栏档案选模型。
中文摘要 AI 辅助
大型语言模型(LLM)在用户反驳后,往往会放弃正确答案,或赞同用户的立场。这种行为被称为谄媚(sycophancy),通常以每个模型的单一比率报告,这几乎无法说明其发生的情境或用户如何避免。我们通过103,939条分级回复研究了产生该行为的条件,这些回复来自十种配置:八个禁用推理的LLM,以及其中两个启用最大推理的配置,所有配置面对相同的200个条目、13种压力条件和四轮对话,每条回复由两个独立的LLM裁判标注。我们发现,主导因素是模型验证用户主张的成本,以及训练好的护栏是否覆盖该主张。从逻辑模型中移除这一任务因素会损失0.485的McFadden R²,而移除模型家族因素损失0.139,移除压力策略因素损失0.009。锚定事实几乎从不被让步(1.3%),而逻辑谜题上的采纳率随着反驳所推答案所需的线索数量增加而上升。个人选择在77.0%的对话中被赞同。困难条目上的大多数让步来自无法可靠解决这些条目的模型;能够解决这些条目的模型很少放弃答案。对于测试的两个模型,最大推理完全消除了这些让步:深度谜题上的采纳率从19.2%和12.5%降至0%。谬误或情绪化框架相比单纯重复没有增加任何效果。三位人工标注者与裁判共识在118/120个校准条目上一致。这些结果为可靠使用提供了实用规则:简化难以验证的问题并深入推理,陈述问题而非个人偏好的答案,对开放性问题要求证据,并根据测量的护栏档案选择模型。
英文摘要
Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it. We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges. We find that the dominant factors are how costly it is for the model to verify the user's claim, and whether a trained guardrail covers it. Removing this task factor from a logistic model costs 0.485 of McFadden $R^2$, against 0.139 for model family and 0.009 for pressure tactic. Anchored facts are almost never conceded (1.3%), while adoption on logic puzzles rises with the number of clues needed to refute the pushed answer. Personal choices are endorsed in 77.0% of conversations. Most concessions on hard items come from models that cannot reliably solve them; models that can solve them rarely give the answer up. For both models tested, maximum reasoning removes these concessions completely: adoption on deep puzzles falls from 19.2% and 12.5% to 0%. Fallacious or emotional framing adds nothing beyond plain repetition. Three human annotators agree with the judges' consensus on 118/120 calibration items. These results give practical rules for reliable use: simplify hard-to-verify problems and reason deeply, state the question rather than one's preferred answer, ask for evidence on open questions, and choose models by their measured guardrail profile.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。