StructReward:用于自修正多模态推理的高效结构化过程奖励
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning
浏览论文内容
中文总结 AI 辅助
本研究提出StructReward框架,通过结构化步骤级奖励对齐结合GRPO目标等方式,在无额外验证器的情况下降低多模态强化学习开销,实现自修正多模态推理。
中文摘要 AI 辅助
基于可验证奖励的强化学习(RLVR)已成为改进多模态推理的有效方法。然而,多数现有方法仅根据最终答案的正确性,用二元奖励评估整个响应,从而丢弃了中间推理步骤中可用的监督信号。过程奖励模型能提供更细粒度的反馈,但通常依赖单独训练的验证器、成本高昂的思维链标注,或大型语言模型(LLM)的在线评判。本研究提出StructReward,这是一个计算高效的框架,通过结构化步骤级奖励对齐提供密集强化信号。StructReward将每个生成的解决方案表示为一系列推理步骤,并使用轻量级数值、符号和词汇匹配规则将其与带过程标签的参考步骤对齐。对齐后的标签被聚合成密集过程奖励,并通过门控组相对策略优化(GRPO)目标与最终答案一致性及输出有效性奖励相结合。我们进一步将策略回滚回收为响应比较和反射自修正的补充监督,而非在策略更新后丢弃它们。此外,我们使用强大的LLM将采样的正确轨迹重写为面向反射的训练实例,进一步增强策略评估和优化其推理的能力。由于奖励计算在线执行,无需额外学习的验证器或外部LLM评判,StructReward大幅降低了多模态强化学习的计算开销。实验结果表明,结构化过程监督和回滚回收为自改进多模态推理提供了一条高效路径。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, thereby discarding the supervision available in intermediate reasoning steps. Process reward models offer finer-grained feedback, but they typically rely on separately trained verifiers, costly chain-of-thought annotations, or online judging by large language models (LLMs). In this work, we introduce StructReward, a compute-efficient framework that provides dense reinforcement signals through structured step-level reward alignment. StructReward represents each generated solution as a sequence of reasoning steps and aligns them with process-labeled reference steps using lightweight numerical, symbolic, and lexical matching rules. The aligned labels are aggregated into a dense process reward and combined with final-answer consistency and output-validity rewards through a gated Group Relative Policy Optimization (GRPO) objective. We further recycle policy rollouts into complementary supervision for response comparison and reflective self-correction, rather than discarding them after policy updates. Separately, we use a strong LLM to rewrite sampled correct trajectories into reflection-oriented training instances, further strengthening the policy's ability to evaluate and refine its reasoning. Since reward computation is performed online without an additional learned verifier or external LLM judge, StructReward substantially reduces the computational overhead of multimodal reinforcement learning. Experimental results show that structured process supervision and rollout recycling provide an efficient path toward self-improving multimodal reasoning.