arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15388cs.AI

精确但未耦合:评审者的精确性并不能保证在多智能体数学推理中采纳批评意见

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur

首次发表
浏览论文内容

中文总结 AI 辅助

研究多智能体数学推理中评审者精确性与采纳批评意见的关系,通过对4181个问题测试发现,仅评审者精确性无法解释结果,广播式讨论有时效果更好,还探究了不同干预对结果的影响,指出以评审者为中心的评估可能高估系统质量。

中文摘要 AI 辅助

许多面向数学和科学的智能体系统采用具有专门评审者角色的分层设计,认为专门的评审阶段应有助于将错误候选答案转变为正确答案。我们使用匹配的gpt-oss-120b智能体对4181个基于验证器的全数学问题进行了测试。协作在最简单层级作用不大,但从第4层级起收益显著增加;在更难的情况下,广播式同行讨论比规划器-执行器-评审器管道(PER)达到更高的最终准确率。我们探究这种差距是由评审者质量还是由批评意见是否改变协议推进的下一个答案来解释。结果表明仅评审者精确性无法解释:PER的评审者比广播式的更精确(0.861对0.644),但评估者验证的有用批评意见改变下一个候选答案的可能性小得多,且产生的评审者引导修复更低。这些结果表明评审者检测质量和采纳批评意见在经验上是可分离的。在匹配的PER干预中,强制明确确认会降低最终准确率,而将评审者指导直接嵌入求解器工作环境中部分改善了后续跟进但未消除差距。总体而言,以评审者为中心的评估可能高估系统质量:一个协议可能能很好地发现错误,但如果不根据这些批评采取行动,仍可能无法解决更多问题。

英文摘要

Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors. Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion reaches higher final accuracy than a planner-executor-reviewer pipeline (PER). We ask whether this gap is explained by reviewer quality or by whether critique changes the next answer the protocol carries forward. It is not explained by reviewer precision alone: PER's reviewer is more precise than broadcast's (0.861 vs. 0.644), yet evaluator-verified useful critique is much less likely to change the next candidate and produces lower reviewer-guided repair. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy, while embedding reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Overall, reviewer-centric evaluation can overstate system quality: a protocol may spot errors well yet still fail to solve more problems if it does not act on those critiques.

发表机构

  • Argonne National Laboratory(阿贡国家实验室)
  • Oregon State University(俄勒冈州立大学)
  • University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑