发表机构
Amazon AGI(亚马逊AGI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Reasoning Jury系统,以多LLM评审团及调控共识机制替代单模型评审员,其识别推理缺陷的准确率显著优于前沿模型,且开销仅为后者的8%-15%,还可用于理解推理LLM的失败模式。
AI 中文摘要
改进推理类大语言模型(LLM)需要具备判断长推理轨迹质量的能力,以实现高效的推理数据整理、强化学习过程中的强训练信号,以及在模型性能评估时深入理解推理行为。此外,识别模型产生的推理错误,可通过运行时提供反馈来提升模型性能。由于长推理轨迹上的该复杂任务难度较高,单模型评审员(即便为前沿模型)在识别推理缺陷方面表现不佳。同时,推理类LLM的在线训练过程中,受使用限制的约束,通常无法使用前沿模型。本研究中,我们提出Reasoning Jury(推理评审团)系统,该系统用由多个LLM组成的评审团及经调控的共识机制替代单评审员,以提升识别推理缺陷的判断保真度。在推理评审团中,通过一次研讨过程揭示推理轨迹的缺陷及其严重程度:一名调控员组织评审团成员开展讨论,成员们互相批判彼此的判断并可修改初始投票;调控员通过评审团间的研讨或判断整合得出共识。我们的研究表明,使用由开放权重模型(如gpt-oss-120b)组成的评审团的Reasoning Jury,在正确识别推理缺陷方面的表现显著优于前沿模型(opus-4.6、sonnet-4.6及gemini-3.1-pro)。除了准确率的提升,评审团的总开销(初始裁决、研讨、整合等)仅为LLM-as-a-judge设置下运行前沿模型开销的一小部分(8%至15%)。我们还展示了如何利用这些判断来理解推理类LLM在基准测试上的失败模式,这能让我们更深入地理解模型的性能。
英文摘要
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation. Additionally, surfacing reasoning mistakes that the model makes would enable improving the model's performance at runtime through providing feedback. Due to the difficulty of this complex task on long reasoning traces, single-model judges (even frontier models) do not do well at identifying reasoning defects. Additionally, leveraging frontier models during online training of reasoning LLMs is generally prohibited due to guardrails in terms of use. In this work, we introduce Reasoning Jury, a system that replaces the single judge with a jury of LLMs and a moderated consensus mechanism, to improve the fidelity of judgments for identifying reasoning defects. In reasoning jury, defects of a reasoning trace and their severity are surfaced through a deliberation where a moderator conducts a discussion amongst the jury where the jurors critique each other's judgments and get to modify their initial votes. The moderator derives a consensus through deliberation amongst jurors or consolidation of judgements. We show that Reasoning Jury with a jury of open-weight models (e.g., gpt-oss-120b) is able to significantly outperform frontier models (opus-4.6, sonnet-4.6, and gemini-3.1-pro) at correctly identifying reasoning defects. Besides accuracy performance improvements, the aggregated cost of the jury (initial verdicts, deliberations, consolidation, etc.) is a fraction (8 to 15%) of the cost of running frontier models in LLM-as-a-judge setup. We also show how these judgements can be leveraged to understand failure modes of reasoning LLMs on benchmarks, which allows much deeper understanding of a model's performance.