发表机构
Aarhus University(奥胡斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型的从众效应,提出抵抗-接受双维度评估,发现多数缓解措施存在此消彼长的边界,仅Reasoning可同时提升抵抗性与接受性。
AI 中文摘要
语言模型的最新进展已实现协作场景,其中多个模型可利用彼此的能力,迭代地改进、转换和扩展彼此的输出。每个智能体在回答前都会看到其他智能体的断言,因此同伴意见会与模型自身的参数知识形成竞争,错误的多数意见可能会推翻模型原本正确的答案。我们在23个开放权重模型、19种条件和3个数据集上测量了这种偏差,共获得超过100万个分级响应。一致的错误多数意见会使模型正确的MMLU答案中22.8%被反转,GPQA上为54.8%,SimpleQA上为71.0%,且84%-89%的反转答案与同伴的答案一致。现有的缓解措施旨在提高抵抗性(即模型在这种压力下保留正确答案的比率),而这仅为协作智能体所需能力的一半。我们将其与接受性(即模型在最初回答错误后采纳同伴正确答案的比率)相结合。我们在这两个维度上对六种方法进行了评分,其中四种来自现有研究,两种为我们自行提出。每种方法仅通过牺牲接受性来获得抵抗性,它们的均值落在单一的抵抗-接受边界上,决定系数R²在0.80至0.90之间。Reflection是已发表的最强方法,它使MMLU抵抗性提高了7.9个点,但接受性下降了15.3个点。Reasoning是唯一的例外:在GPQA和SimpleQA上,它的权衡与其他方法类似;但在模型可自行推导答案的MMLU主题上,它同时使抵抗性提高7.2个点、接受性提高9.6个点,是我们发现的唯一一种能同时提升这两个指标的干预措施。
英文摘要
Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with $R^2$ between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.