发表机构
University of Pennsylvania; DeepKeep(宾夕法尼亚大学; DeepKeep(注:根据上下文,此处为作者曾任职的机构,需保留英文名称))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过谄媚滞后指标评估模型在用户施压后恢复的能力,发现重置不能消除偏差,而移除压力历史或添加可信证据可显著提升恢复效果。
AI 中文摘要
基于事实的语言模型通常通过添加相关上下文来评估,但多轮对话中也包含可能污染后续事实回答的未经证实的用户主张。我们研究了压力后恢复能力:即当用户反复主张一个错误答案并随后撤回该压力时,模型是否能够恢复到干净上下文行为。我们引入了一种用于多项选择事实对话的压力后恢复协议,并测量了谄媚滞后,即相对于干净上下文反事实,模型对用户主张的错误答案所保留的残余概率。在七个指令微调的开源模型和两个事实基准上,普通的重置通常能减少但并不能消除压力引起的偏差。保留历史的修复措施,如用户撤回、系统重置和自我验证,在严格的干净恢复诊断下仅恢复了14个模型-数据集对中的2-3个,而改变有效上下文的操作则明显更可靠;两种完全移除承载压力的历史的条件,即全新上下文删除和上下文截断,恢复了14/14。在跨越十四个模型-数据集对的预言可信证据条件下,保留承载压力的历史同时添加基准派生的可信证据,将准确率从0.368提高到0.929,而错误答案跟随率从41.2%下降到4.3%。对照实验表明,该效应不能由对话长度、重复的信心、合理的干扰项、仅仅提及错误答案或选项标签惯性来解释。这些结果表明,忠实的基于事实的对话要求评估哪些先前的上下文应被视为证据,哪些应在回答前被移除或隔离。
英文摘要
Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocates a wrong answer and then withdraws that pressure. We introduce a recovery-after-pressure protocol for multiple-choice factual dialogue and measure sycophancy hysteresis, the residual probability assigned to the user-advocated wrong answer relative to a clean-context counterfactual. Across seven instruction-tuned open-weight models and two factual benchmarks, ordinary reset often reduces but does not erase pressure-induced bias. History preserving repairs such as user retraction, system reset, and self-verification recover only 2-3/14 model-dataset pairs under the strict clean-restoration diagnostic, whereas operations that change the effective context are substantially more reliable; the two conditions that remove the pressure-bearing history entirely, fresh-context deletion and context truncation, recover 14/14. In an oracle trusted-evidence condition across fourteen model-dataset pairs, preserving the pressure-bearing history while adding benchmark-derived trusted evidence increases accuracy from 0.368 to 0.929, while wrong-answer following falls from 41.2% to 4.3%. Controls show that the effect is not explained by dialogue length, repeated confidence, plausible distractors, mere false-answer mention, or option-label inertia. These results suggest that faithful grounded dialogue requires evaluating which prior context should be treated as evidence and which should be removed or quarantined before answering.