arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当上游消息覆盖正确答案:多智能体LLM协作的受控研究

When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen, Chenxiao Fan, Jiakai Wang, Dong Wang, Yang Liu, Wenjie Wang, Xiangnan He

arXiv 2609.36855首次发表:更新:

发表机构

University of Science and Technology of China; Qwen Business Unit of Alibaba; National University of Singapore(中国科学技术大学; 阿里巴巴Qwen业务部; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验揭示多智能体LLM协作中,上游错误消息可导致下游智能体覆盖正确答案(高达32%案例),并识别出“答案替换”模式,建议基于上游可靠性选择性通信以恢复准确率。

AI 中文摘要

多智能体大语言模型(LLM)系统依赖专门化智能体之间的消息传递来完成复杂任务。然而,上游智能体可能提供有用信息或错误答案,导致下游智能体覆盖其自身证据支持的正确答案。先前的研究未能清晰区分通信带来的益处与错误消息造成的损害。我们通过跨五个基准和五个接收器的受控实验研究该问题,在保持下游任务和证据固定的同时,比较三种条件下的答案:无消息、上游智能体的原始消息、或结论相反的消息。实验揭示三个关键发现。第一,当下游智能体原本会答错时,消息通常有帮助。第二,消息也可能有害:当下游智能体在没有消息时会答对时,错误的上游消息在多达32%的情况下改变了答案。第三,在94%的经审计的有害案例中,下游智能体复制了上游的具体错误答案——我们将此模式称为答案替换。移除不可靠的消息可恢复部分丢失的准确率,这表明通信应基于上游可靠性和下游智能体已有的证据进行选择性进行。

英文摘要

Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑