arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22700cs.CLcs.AI

LLaDA-PRM: 一种双向步骤级推理评估器

LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator

Yiming Feng, Naihao Deng, Yulong Chen, Rada Mihalcea

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出LLaDA-PRM,一种基于双向注意力的8B步骤级推理评估器,通过对比实验验证双向注意力优于因果注意力,并在多个基准上显著超越现有方法,同时可用于训练数据选择。

中文摘要 AI 辅助

步骤级推理评估器通常基于自回归语言模型,其因果注意力机制将每个步骤的表示限制为问题、先前步骤和当前步骤。然而,当完整解决方案可用时,早期步骤的有效性可能只有通过其下游后果才能变得更加清晰。我们通过一项受控的54次运行比较来验证这一假设,比较了1B至3B规模的因果和双向LLaDA评估器,仅改变自注意力掩码,发现双向注意力带来了一致的改进。基于这一发现,我们引入了\prm{},一个8B的双向评估器,在MR-MATH-invalid上达到了88.8的步骤级F1分数,在分布外的MR-GSM8K原始问题子集上达到了83.8,分别比ReasonEval-Llemma-34B高出11.3和10.3个F1点。\prm{}在在线设置中评估不完整推理轨迹时也保持有效,在两个基准上大幅超越了最强的基线。我们进一步表明,\prm{}提供了有效的训练数据选择信号,提升了Mistral-7B在MATH-500上的性能。

英文摘要

Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal attention restricts each step representation to the problem, previous steps, and the current step. Yet, when the complete solution is available, the validity of an earlier step may become clearer only through its downstream consequences. We validate this hypothesis through a controlled 54-run comparison of causal and bidirectional LLaDA evaluators at 1B--3B scale, changing only the self-attention mask, and find bidirectional attention yields consistent improvements. Building on this finding, we introduce \prm{}, an 8B bidirectional evaluator that reaches 88.8 step-level F1 on MR-MATH-invalid and 83.8 on the out-of-distribution MR-GSM8K original-question subset, outperforming ReasonEval-Llemma-34B by 11.3 and 10.3 F1 points, respectively. \prm{} also remains effective when evaluating incomplete reasoning traces in online settings, outperforming the strongest baselines on both benchmarks by a large margin. We further show that \prm{} provides an effective training-data selection signal, improving Mistral-7B performance on MATH-500.

发表机构

  • University of Michigan(密歇根大学)
  • University of Aberdeen(阿伯丁大学)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑