arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重置并非恢复:通过谄媚滞后评估从虚假对话上下文中恢复的能力

Reset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy Hysteresis

Adi Shnaidman

arXiv 2609.33672首次发表:更新:

发表机构

University of Pennsylvania; DeepKeep(宾夕法尼亚大学; DeepKeep(注:根据上下文,此处为作者曾任职的机构,需保留英文名称))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过谄媚滞后指标评估模型在用户施压后恢复的能力,发现重置不能消除偏差,而移除压力历史或添加可信证据可显著提升恢复效果。

AI 中文摘要

基于事实的语言模型通常通过添加相关上下文来评估,但多轮对话中也包含可能污染后续事实回答的未经证实的用户主张。我们研究了压力后恢复能力:即当用户反复主张一个错误答案并随后撤回该压力时,模型是否能够恢复到干净上下文行为。我们引入了一种用于多项选择事实对话的压力后恢复协议,并测量了谄媚滞后,即相对于干净上下文反事实,模型对用户主张的错误答案所保留的残余概率。在七个指令微调的开源模型和两个事实基准上,普通的重置通常能减少但并不能消除压力引起的偏差。保留历史的修复措施,如用户撤回、系统重置和自我验证,在严格的干净恢复诊断下仅恢复了14个模型-数据集对中的2-3个,而改变有效上下文的操作则明显更可靠;两种完全移除承载压力的历史的条件,即全新上下文删除和上下文截断,恢复了14/14。在跨越十四个模型-数据集对的预言可信证据条件下,保留承载压力的历史同时添加基准派生的可信证据,将准确率从0.368提高到0.929,而错误答案跟随率从41.2%下降到4.3%。对照实验表明,该效应不能由对话长度、重复的信心、合理的干扰项、仅仅提及错误答案或选项标签惯性来解释。这些结果表明,忠实的基于事实的对话要求评估哪些先前的上下文应被视为证据,哪些应在回答前被移除或隔离。

英文摘要

Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocates a wrong answer and then withdraws that pressure. We introduce a recovery-after-pressure protocol for multiple-choice factual dialogue and measure sycophancy hysteresis, the residual probability assigned to the user-advocated wrong answer relative to a clean-context counterfactual. Across seven instruction-tuned open-weight models and two factual benchmarks, ordinary reset often reduces but does not erase pressure-induced bias. History preserving repairs such as user retraction, system reset, and self-verification recover only 2-3/14 model-dataset pairs under the strict clean-restoration diagnostic, whereas operations that change the effective context are substantially more reliable; the two conditions that remove the pressure-bearing history entirely, fresh-context deletion and context truncation, recover 14/14. In an oracle trusted-evidence condition across fourteen model-dataset pairs, preserving the pressure-bearing history while adding benchmark-derived trusted evidence increases accuracy from 0.368 to 0.929, while wrong-answer following falls from 41.2% to 4.3%. Controls show that the effect is not explained by dialogue length, repeated confidence, plausible distractors, mere false-answer mention, or option-label inertia. These results suggest that faithful grounded dialogue requires evaluating which prior context should be treated as evidence and which should be removed or quarantined before answering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑