arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04270cs.SEcs.CLcs.MA

评审者能力决定拒绝目标而非修复技能:来自大语言模型执行-评审-修正流程的证据

Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines

Faizan Tanveer

首次发表
浏览论文内容

中文总结 AI 辅助

该研究在100道奥林匹克数学题上,探究LLM执行-评审-修正流程中评审者能力的影响,发现中端跨系列评审者可提升准确率,自评审虽错误检测率高但增益有限,最弱评审者无作用。

中文摘要 AI 辅助

多智能体大语言模型(LLM)流程日益将执行、验证等角色分配给不同能力层级的模型,原因是在每个阶段都使用旗舰模型成本过高。现有文献已证实验证阶段并非总是有益,但研究中评审者能力相对于执行者大致固定。本研究改变了这一设定:将评审者替换为能力范围低至完全无法解决问题的模型,在固定的100道奥林匹克数学题集上,测量每一次单独拒绝的结果。跨系列的中端评审者使最终准确率提升12个百分点,从52%升至64%(p=0.0005),且零错误答案受损。同模型自评审达到所有条件中最高的错误检测率(召回率0.85),但未产生显著增益:其拒绝频率为跨系列评审者的2.1倍,修复率仅为后者的三分之一(15%对43%,p=0.0074),且将自身正确答案错误拒绝的比例为35%,而跨系列评审者仅为2%(配对p=0.000015)。自评审的低损伤率被证实是修正惯性的产物,而非评审者质量:在18次被错误拒绝的正确答案中,执行者遵从的3次均变为错误,而被忽略的15次则保持不变。低于能力下限后,评审者角色变得无效:本研究中最弱的评审者未改变100道题的任何最终答案,同时将令牌成本翻倍。这些发现描述了100道题上的单一执行者-评审者配置,应视为受控试点而非关于验证阶段的一般性结论。

英文摘要

Multi-agent LLM pipelines increasingly assign roles, including execution and verification, to models of different capability tiers. This is done because running a flagship model at every stage is expensive. Previous literature has established that verification stages are not always beneficial, but holds reviewer capability roughly fixed relative to the executor. We vary it. We replace the reviewer with models spanning a capability range down to one that cannot solve the problems at all, and measure the outcome of every individual rejection. This is done across a constant set of 100 olympiad mathematics problems. A cross-family mid-tier reviewer improves final accuracy by 12 percentage points, from 52 to 64 percent (p = 0.0005), with zero damaged answers. Same-model self-review attains the highest error-detection rate of any condition (0.85 recall) yet yields no significant gain: it rejects 2.1 times as often for a third the repair rate (15 against 43 percent, p = 0.0074) and falsely rejects 35 percent of its own correct answers against 2 percent for the cross-family reviewer (paired p = 0.000015). The low damage rate of self-review proves to be an artifact of revision inertia rather than reviewer quality: of 18 falsely rejected correct answers, the three where the executor complied all became wrong, while the fifteen it ignored survived unchanged. Below a capability floor the role becomes inert: our weakest reviewer changed zero of 100 final answers while doubling token cost. These findings describe a single executor-reviewer configuration on 100 problems and should be read as a controlled pilot rather than a general claim about verification stages.

发表机构

  • National University of Computer and Emerging Sciences (FAST NUCES)(国立计算机与新兴科学大学(FAST NUCES))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑