arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

沉默的异议:屈服于多数派的LLM智能体仍代表其原始前提

Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise

Ziang Ni, Peng Zou

arXiv 2610.02702首次发表:更新:

发表机构

Delft University of Technology; Sun Yat-sen University(代尔夫特理工大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过两跳事实问题探究LLM智能体在多数压力下改变答案时是否仍保留原始前提,使用雅可比透镜读取内部表示,发现屈服智能体在内部仍代表原始桥梁,表明陈述共识可能夸大真实一致性。

AI 中文摘要

多智能体辩论日益被用于在LLM智能体之间达成共识,然而智能体常常屈服于一致多数。当一个智能体改变其答案时,它是改变了想法还是仅仅改变了陈述?我们通过两跳事实性问题来研究这一点,其中间实体(即桥梁,例如“圣家堂所在国家的首都”中的国家)从未被任何人提及。脚本化的同伴,扮演阿希(Asch)式同盟者的角色,一致断言一个取自另一个具有不同桥梁的事实的错误答案。在智能体回答的那一刻,我们使用雅可比透镜(J-lens)从其残差流中读取桥梁,并为了比较,也使用logit透镜。在针对四个开放权重模型的预注册测试中,Qwen3.5-4B、Qwen3.6-27B和Gemma-4-E4B-it的智能体尽管屈服,但在输出层以下的预注册层中仍代表其原始桥梁(相对于对照实体的hit@100分别为0.85、0.22和0.24),而logit透镜很少将其排在前100个token中(0.00-0.06)。这些智能体也代表了同伴答案背后的桥梁,超出了提及基线。一个预注册的附录隐藏了智能体较早的答案或将其移除:屈服的智能体在所有四个模型中仍代表其原始桥梁(隐藏答案时分别为0.43、0.29、0.37和0.25),包括Llama-3.1-8B-Instruct,而在其答案可见时,该模型几乎不这样做(0.03)。因此,前提可以仅从问题中计算出来,而智能体陈述的是多数的答案。隐藏较早的答案也改变了从众行为:Qwen3.5-4B在89%的问题上屈服,而不是8%。在探索性干预中,注入桥梁的J-lens方向仅在两个Qwen模型中使智能体回到其原始答案。因此,多智能体辩论中的陈述共识可能夸大了一致性。我们还报告了我们预注册程序的负面结果。

英文摘要

Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in "the capital of the country where the Sagrada Familia is located") is never stated by anyone. Scripted peers, in the role of Asch's confederates, unanimously assert a wrong answer taken from another fact with a different bridge. At the moment the agent answers, we read the bridge from its residual stream with the Jacobian lens (J-lens) and, for comparison, the logit lens. In pre-registered tests on held-out facts with four open-weight models, agents of Qwen3.5-4B, Qwen3.6-27B and Gemma-4-E4B-it that gave in still represented their original bridge in the pre-registered layers below the output (hit@100 above a control entity: 0.85, 0.22 and 0.24), where the logit lens rarely ranked it among the top 100 tokens (0.00-0.06). These agents also represented the bridge behind the peers' answer, beyond a mention baseline. A pre-registered addendum hid the agent's earlier answer or removed it: agents that gave in still represented their original bridge in all four models (0.43, 0.29, 0.37 and 0.25 with the answer hidden), including Llama-3.1-8B-Instruct, which barely did so with its answer in view (0.03). The premise can thus be computed from the question alone while the agent states the majority's answer. Hiding the earlier answer also changed conformity: Qwen3.5-4B gave in on 89% of questions instead of 8%. In exploratory interventions, injecting the bridge's J-lens direction brought agents back to their original answer only in the two Qwen models. Stated consensus in multi-agent debate can thus overstate agreement. We also report the negative results of our pre-registered program.

Comments9 pages, 4 figures, 3 tables. Supplementary material in ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑