发表机构
Illinois Institute of Technology(伊利诺伊理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示多智能体LLM系统中过滤有害同伴一致性受限于模型自我认知,提出“墙”与“悬崖”现象,强调修订前信息优于事后过滤。
AI 中文摘要
多智能体大语言模型系统预期会更加可靠,因为智能体可以相互发现错误。但同伴压力是双向的:纠正错误答案的同一修正也可能推翻正确答案。诱人的保障措施是设置一个刹车,以保留有益的修订并阻止有害的修订。我们表明,这个刹车很难构建,原因很简单:当且仅当原始答案正确时,修订才是有害的,因此决定是否阻止它等同于知道模型是否已经正确。这将寻找刹车的开放式任务转化为一个可测量的量,即模型的自我认知:任何基于部署时信号构建的刹车实际上都是一个正确性探针,而自我认知远非完美(六个模型家族的AUROC约为0.64至0.89)。我们将这一上限称为“墙”。即使是白盒引导模型自身的正确性方向也无法突破它:它改变了模型修订的频率,但有害和有益的修订会同时变化。在群体规模上,这堵墙变成了悬崖:当大多数智能体初始出错时,辩论会将共同的错误放大为自信的错误共识。在我们的多项选择社会中,更多的智能体、更大的模型多样性和更强的成员都无法解决这个问题。有帮助的是在修订前添加信息,而不是在修订后过滤。局部一致并非全局正确。
英文摘要
Multi-agent LLM systems are expected to be more reliable because agents can catch each other's mistakes. But peer pressure cuts both ways: the same correction that fixes a wrong answer can overturn a right one. The tempting safeguard is a brake that keeps the beneficial revisions and blocks the harmful ones. We show this brake is hard to build, for a simple reason: a revision is harmful exactly when the original answer was right, so deciding whether to block it is the same as knowing whether the model was already correct. This turns the open-ended hunt for a brake into one measurable quantity, the model's self-knowledge: any brake built from a deploy-time signal is a correctness probe in disguise, and self-knowledge is far from perfect (AUROC $\approx 0.64$--$0.89$ across six model families). We call this ceiling the wall. Even white-box steering of the model's own correctness direction does not breach it: it changes how often the model revises, but harmful and beneficial revisions move together. At population scale the wall becomes the cliff: when most agents start wrong, debate amplifies the shared mistake into a confident, wrong consensus. In our multiple-choice societies, more agents, more model diversity, and a stronger member do not fix it. What helps is adding information before the revision, not filtering after it. Local agreement is not global correctness.