AI 中文总结
本研究审查谄媚基准SycEval,发现其模板中两个混淆变量(目标答案命名和输出格式指令位置)导致时机效应结论错误,并提出了可纠正的检查方法。
AI 中文摘要
谄媚是指语言模型在用户反驳时屈服的倾向,即放弃正确答案而迎合用户的说法。多个基准通过编写反对意见并记录模型屈服的频率来测量这一现象。由于反对意见是一个提示模板,模板中其他变化的内容会与被声称隔离的属性一同被测量。我们审查了SycEval,该基准报告称在模型回答之前提出的反对意见(预先)比在回答之后提出的反对意见(上下文中)导致更多屈服,并将这一差异归因于时机。其模板的两个特征与被测属性同时变化。首先,SycEval的反对意见通过四个强度级别升级,在最弱的两个级别中,只有预先模板指定了目标答案,因此时机和命名同时变化。我们构建了缺失的比较,并在多项选择题和SycEval自身的自由形式流程上进行了测试。指定目标答案将跟随率(即与用户断言匹配的样本比例)提高了14.1至49.5个百分点(pp),而一旦两个模板都指定了目标答案,时机比较在我们测试的五个模型条件中的三个上发生了逆转:模型在预先反对下的屈服程度低于上下文反对,与SycEval相反。其次,基准为自动评分附加的输出格式指令也随反对意见的位置而变化。将其从反驳中移到问题中,会在相同项目上使不同模型的效果朝相反方向翻转(p=0.0059)。同样的效果出现在SycEval的自由形式流程中:将其指令重新定位使Llama-3.1-8B的屈服率降低了5.1个百分点,而等价检验确认对Qwen3-4B没有影响。两个混淆变量都存在于模板中而非被测模型中,因此两者都是可纠正的:我们最后提出了基准作者在发布前可以应用的三个检查。
英文摘要
Sycophancy is a language model's tendency to cave when a user pushes back, abandoning a correct answer for the user's. Several benchmarks now measure it by scripting an objection and recording how often the model caves. Because that objection is a prompt template, whatever else the template varies is measured along with the property it claims to isolate. We audit SycEval, which reports that objections raised before a model answers (preemptive) cause more caving than those raised after (in-context), and attributes the gap to timing. Two features of its templates vary alongside the property each is meant to test. First, SycEval's objections escalate through four strength levels, and at the two weakest only the preemptive template names a target answer, so timing and naming vary together. We build the missing comparison and test it on multiple-choice questions and SycEval's own free-form pipeline. Naming a target answer raises the follow rate, the share of samples matching the user's assertion, by 14.1--49.5 percentage points (pp), and once both templates name one, the timing comparison reverses on three of the five model conditions we test: models cave \emph{less} under preemptive objections than in-context ones, opposite to SycEval. Second, the output-format instruction benchmarks append for automatic grading also varies with objection placement. Moving it from the pushback into the question flips its effect in opposite directions across models on the same items ($p=0.0059$). The same effect appears in SycEval's free-form pipeline: relocating its instruction lowers caving by 5.1pp on Llama-3.1-8B, while an equivalence test confirms no effect on Qwen3-4B. Both confounds live in the template rather than the models under test, so both are correctable: we close with three checks benchmark authors can apply before publishing.
Comments22 pages, 19 tables, 3 figures