arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

错误但有用:多智能体消息中超越答案正确性的轨迹价值

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur

arXiv 2608.14375首次发表:更新:

AI 中文总结

该研究提出DHD协议,发现多智能体错误消息含有用信息,错误但有帮助的消息普遍存在,其轨迹价值可辅助决策,答案正确性不决定轨迹价值。

AI 中文摘要

多智能体推理系统常使用一致性、置信度或自动评分来决定哪些消息应影响最终答案,这类过滤假设大概率正确的消息也值得保留。然而错误答案可能包含有用的分解、约束或科学原理。我们通过Diverse Hypothesis Deliberation(DHD,一种受控测量协议)测试这种区分:该协议缓存5条独立生成的消息,在每条消息可用或隐藏的情况下,重新运行名为integrator的下游求解器,通过重放比较衡量消息的轨迹价值,即该消息的可用性是否有助于或损害后续推理。在5个数学和科学基准及两个公开模型家族gpt-oss-120b和gemma-4-31B-it中,错误但有帮助的消息出现在每一个基准-模型组合中。在改变最终正确性的错误答案消息中,每个模型里超过十分之四的改变是有帮助的。受控重复实验显示,可重复的消息效应数量不太可能仅由重放变异导致(p=0.0002)。对可重复的错误-有帮助消息的针对性干预发现,完整消息效果最佳,保留其推理过程比仅保留答案能维持更多成功;完整消息优势的来源仍待探究。在同一问题中,重复的轨迹价值证据也能比仅靠答案正确性识别出更好的保留或移除选择。因此,答案正确性具有参考性但不决定轨迹价值,DHD可测量这一缺失属性,并生成可重复使用的标签以学习智能体应何时倾听。

英文摘要

Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test this distinction with Diverse Hypothesis Deliberation (DHD), a controlled measurement protocol that caches five independently generated messages and replays the same downstream solver, called the integrator, with each message available or hidden. The replay comparison measures a message's trajectory value: whether making the message available helps or harms subsequent reasoning. Across five mathematics and science benchmarks and two openly available model families, gpt-oss-120b and gemma-4-31B-it, wrong-helpful messages appear in every benchmark-model combination. Among wrong-answer messages that change final correctness, more than four in ten changes are helpful in each model. Controlled repeats show that the number of repeatable message effects is unlikely to arise from replay variation alone (p=0.0002). A focused intervention on repeatable wrong-helpful messages finds that the complete message works best, while retaining its reasoning preserves more success than retaining only its answer; the source of the complete-message advantage remains open. Within the same problem, repeated trajectory-value evidence also identifies a better keep-or-remove choice than answer correctness alone. Answer correctness is therefore informative but does not determine trajectory value. DHD measures this missing property and produces reusable labels for learning when agents should listen.

Comments24 pages, 9 figures. Includes an appendix and an ancillary reproducibility artifact

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑