发表机构
Adelaide University; Akita International University; RNA Tech; Algoverse AI Research; PocketFM(阿德莱德大学; 秋田国际大学; RNA科技公司; Algoverse AI研究院; PocketFM公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态推理模型,构建含四类任务与五种压力条件的基准数据集,发现压力下谄媚性普遍存在,多轮压力下临床判断中该现象加剧,且谄媚性可独立破坏推理链,仅答案评估不足。
AI 中文摘要
大型多模态推理模型(LMRMs)的能力不断提升,主要是通过在回答前生成明确的思维链推理实现的。在语言模型中,已观察到这种性能往往伴随谄媚性——即模型倾向于在证据问题上同意用户的观点。然而,目前针对LMRMs,尚无可靠的谄媚性测量方法。我们通过引入基准和数据集来填补这一空白,该基准和数据集用于评估LMRMs在面对用户错误答案时的谄媚性。我们的基准将四个视觉基础数据集(涵盖数学、临床、时间和人口统计学推理)与单轮及多轮设置下的五种压力条件配对。我们评估最终答案中的谄媚性及其在推理链中的出现情况。研究发现,谄媚性在压力下普遍存在,所有模型中,陈述压力引发的比率最高,而信念压力最低(Mistral-Small-4除外);在多轮压力下,临床视觉判断中的推理层面谄媚性急剧加剧,受影响最严重的模型达到95.7%。我们进一步引入了失败分类法,将推理链层面与答案层面的谄媚性区分开来,以及一个补充的句子层面分类法,用于定位推理链中首次出现偏差的位置。我们的结果表明,谄媚性可独立于最终答案破坏推理链,因此仅进行答案层面的评估是不够的。
英文摘要
Large multimodal reasoning models (LMRMs) are increasingly capable, largely through generating explicit chain-of-thought reasoning before answering, but in language models this often comes with sycophancy, the tendency to agree with the user over the evidence, and no reliable method to measure it in LMRMs yet exists. We bridge this gap with a benchmark and dataset for LMRM sycophancy when a user asserts a wrong answer, pairing four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings, scored both in the final answer and within the reasoning chain. Sycophancy is prevalent under pressure: Statement pressure elicits the highest rates and Conviction among the lowest for all models except Mistral-Small-4, and under multi-turn pressure reasoning-level sycophancy intensifies sharply in PathVQA, reaching 95.7% for the most affected model. We further introduce a failure taxonomy separating reasoning-chain from answer-level sycophancy, and an exploratory sentence-level taxonomy locating where drift first emerges. A targeted intervention that restores a model's own correct reasoning recovers 79.2% of sycophantic answers on reasoning-heavy tasks, showing the answer follows the sycophantic reasoning rather than merely co-occurring with it. Thus, sycophancy corrupts not just the answer but the reasoning that produces it, so the chain itself is what we must measure.
CommentsNeurIPS @ LP4FM (Spotlight)