发表机构
Honda Research Institute Japan Co., Ltd.(日本本田研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型网络中消息容量与主张措辞如何共同决定集体求真结果,发现二者设定共识转变点,决定正确或错误共识的达成。
AI 中文摘要
无论是人类还是大型语言模型(LLM),讨论中的智能体仅阅读其他参与者贡献中的少数几条,受限于认知、上下文或成本。即使多数智能体初始正确,LLM集体也可能达成错误共识;我们探究仅凭这种阅读限制在多大程度上决定结果。我们用单一数字——消息容量——来建模该限制,它设定智能体阅读其他消息的数量,并据此生成通信网络。在31,824次随机化查询中,我们发现一个80亿参数模型对主张的判断实际上简化为其收件箱加权和的逻辑函数,即具有除法归一化权重的随机二元神经元的更新规则。仅依据这些权重和网络的度统计,一旦智能体平均阅读其31个来源中的少于6.4条,错误共识应从任何起点都不可达。在1,414个指定起点的回合中,预测失败:从每个起点,正确方在少于50%的回合中获胜,且当75%的智能体初始正确时,正确方仅在28-45%的回合中获胜。失败可追溯到场(field),即主张措辞在任何消息被阅读前为智能体答案设定的阈值:实验主张的场低于校准均值,且使用每个主张自身的场,相同权重重现了结果。反转措辞表明,该阈值遵循主张所断言的内容,而非其真实性。在第二个80亿模型上,该流程预测了依赖于主张的双稳态;转变点出现在计算位置,且八主张校准在16个条件中的15个匹配。在700亿规模下未检测到断言偏差。因此,集体的命运主要由两个单智能体测量决定:主张措辞设定的阈值,以及设定转变点的消息容量。
英文摘要
Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others' messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model's judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network's degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim's wording sets for the agent's answer before any message is read: the experimental claims' fields lay below the calibration mean, and with each claim's own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective's fate is largely set by two single-agent measurements: the threshold a claim's wording sets, and the message capacity that sets the transition point.