发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ICR框架,通过固定初始推理并对比无消息修订,审计LLM多智能体通信,发现相似准确率下修订行为差异大,通信质量非通道固有属性。
AI 中文摘要
多智能体通信旨在帮助智能体从彼此的信息中获益。然而,系统性能的提升留下了一个根本性的模糊之处:这些提升反映的是有效的通信、有利的智能体架构,还是仅仅额外的推理?由于通信方法通常在其所设计的系统内进行评估,这些因素难以区分。最终准确率进一步将纠正的错误和破坏的答案合并为一个单一结果,掩盖了通信如何改变决策。我们引入了独立-通信-修订(ICR),一个受控框架,将通信评估为独立推理后的答案修订。ICR固定初始推理轨迹,测量基于两个智能体初始正确性的纠正和保留,并使用无消息修订对照来量化超出额外推理的收益。在四个推理基准上,我们对文本和潜在通信的审计揭示,相似的总体准确率可能掩盖显著不同的修订行为。与仅传输答案相比,完整推理在所有四个基准上增加了纠正同时减少了保留,因此更丰富的消息放大了有益和有害的影响。在MedQA和GPQA-D上的接收者策略比较进一步表明,结构化验证策略将每个通道转向更高的保留和更低的纠正,而其选择性效果在不同通道和任务间有所变化。这些发现挑战了将通信质量视为通道内在属性的观点。因此,ICR将评估重心重新聚焦于选择性修订,提供了一个统一框架来检验消息内容和接收者策略如何共同产生益处和危害。
英文摘要
Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.