发表机构
National University of Singapore; Yunnan University; Agency for Science, Technology and Research (A*STAR), Singapore(新加坡国立大学; 云南大学; 新加坡科技研究局(A*STAR))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
多智能体协作中,即使决策正确也可能留下损坏信息状态;OffQuery揭示此问题,ReGround通过解决冲突、验证事实、重建状态,显著提升证据验证、状态重建和任务解决能力。
AI 中文摘要
多智能体系统通常以其是否达到正确答案来评判。这可能忽略一种独特的失败:即使即时决策是正确的,协作也可能留下损坏的信息状态。我们称之为查询外失败。为了研究协作决策支持中的这种失败,我们引入了OffQuery,它在两个具有代表性的高风险场景中分别评估证据验证(T1)、共享状态重建(T2)和任务解决(T3):医疗保健和灾难响应。在GPT、Gemini和Qwen模型中,标准协作的任务性能远强于状态可靠性。在21个模型-设置组合中平均,任务解决率达到64.7%,而证据验证和状态重建仅分别达到14.3%和43.1%。我们将这一差距追溯到选择性信息使用:当前查询通常绕过损坏的事实,这些事实在后续任务需要时变得至关重要。我们进一步引入了ReGround,它解决冲突证据、验证共享事实、重建可信状态,并基于该状态进行推理。在来自三个家族的七个模型中,ReGround在每种评估设置中都提升了所有三种能力,在T1、T2和T3上的平均相对增益分别为309.0%、82.9%和17.6%。因此,可靠的协作既需要正确的决策,也需要可靠的共享状态以供未来推理。
英文摘要
Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introduce OffQuery, which separately evaluates evidence verification (T1), shared-state reconstruction (T2), and task resolution (T3) in two representative high-stakes settings: healthcare and disaster response. Across GPT, Gemini, and Qwen models, standard collaboration shows much stronger task performance than state reliability. Averaged over 21 model--setting combinations, task resolution reaches 64.7%, while evidence verification and state reconstruction reach only 14.3% and 43.1%. We trace this gap to selective information use: current queries often bypass corrupted facts, which become consequential when later tasks require them. We further introduce ReGround, which resolves conflicting evidence, verifies shared facts, reconstructs a trusted state, and reasons over that state. Across seven models from three families, ReGround improves all three capabilities in every evaluated setting, with average relative gains of 309.0%, 82.9%, and 17.6% on T1, T2, and T3. Reliable collaboration therefore requires both a correct decision and a reliable shared state for future reasoning.