arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化步骤级推理以实现大语言模型的有效自校正

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan

arXiv 2608.11573首次发表:更新:

发表机构

Nanyang Technological University; National University of Singapore; VinUniversity; Institute for Infocomm Research (IR), A*STAR(南洋理工大学; 新加坡国立大学; VinUniversity; 新加坡科技研究局信息通信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型自校正的核心挑战,提出SFS-DPO及教师辅助变体SFS-DPO-R框架,经多模型多域评估,其性能优于现有步骤级训练基线,可提升自校正频率与有效性。

AI 中文摘要

实现让模型验证并纠正自身错误的有效自校正,仍是大语言模型(LLMs)面临的核心挑战。本研究提出Self-Fix Step-DPO(SFS-DPO),这是一种基于强化学习的两阶段框架,用于步骤级自验证与自校正:第一阶段通过步骤级偏好优化强化步骤级推理,第二阶段明确训练模型进行自验证与自校正。我们进一步引入教师辅助变体SFS-DPO-R,其纳入错误验证的解释性理由以提供更强的校正信号。对多个LLMs开展的全面域内及域外评估表明,SFS-DPO与SFS-DPO-R始终优于现有的步骤级训练基线。我们的分析还揭示,模型的自校正频率与有效性均得到提升,凸显了强化步骤级推理对实现鲁棒性能的重要性。

英文摘要

Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self-verification and self-correction. The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct. We further introduce a teacher-assisted variant, SFS-DPO-R, which incorporates explanatory rationales for error verification to provide stronger corrective signals. Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines. Our analysis further reveals improvements in self-correction frequency and effectiveness, highlighting the importance of strengthening step-level reasoning for robust performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑