arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

下一块推理强化学习(RL)真的比监督微调(SFT)更好?——在无思维链(CoT)数据下重新审视训练策略

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu, Ziyi Wang, Xun Zhao, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen

arXiv 2608.23256首次发表:更新:

发表机构

University of Science and Technology of China; Shanghai AI Laboratory(中国科学技术大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比下一块推理RL与混合SFT,发现混合SFT在更少计算量下性能上限更高,且需在完整后训练 pipeline 中评估无CoT训练策略。

AI 中文摘要

近期研究提出了下一块推理强化学习(RL),用于利用无思维链(CoT)数据——这类数据包含解题过程、教材推导等富含推理内容但缺乏显式思维链标注的语料。该方法训练模型生成隐式推理轨迹,并通过预测下一块文本的能力对其进行奖励。尽管该方法颇具前景,但现有评估主要将其与传统SFT基线进行比较,仍存在疑问:性能提升是源于RL公式本身,还是源于更有效地让模型接触无CoT数据?本研究针对该问题开展受控实验,对比下一块推理RL与一种简单但此前被忽视的替代方案:混合SFT,即仅在一个监督微调阶段中同时使用无CoT数据和长CoT数据进行训练。尽管混合SFT结构简单,但其在RLVR后的性能上限明显高于下一块推理RL,且所需训练计算量减少超过60倍。该优势在领域内数学推理和领域外推理任务中均保持一致。此外,研究还发现,RLVR前的准确率更高并不一定能转化为RLVR后的准确率更高,强调需在完整的后训练 pipeline 背景下评估无CoT训练策略。

英文摘要

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method trains a model to generate implicit reasoning traces and rewards them by their ability to predict the next chunk of text. While promising, existing evaluations primarily compare against conventional SFT baselines, leaving open whether the gains come from the RL formulation itself or from more effectively exposing the model to no-CoT data. We address this question with a controlled study of next-chunk reasoning RL and a simple but previously overlooked alternative: Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data. Despite its simplicity, Mixed SFT achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute. The advantage is consistent across in-domain mathematical reasoning and out-of-domain reasoning tasks. Moreover, we show that higher pre-RLVR accuracy does not necessarily translate into higher post-RLVR accuracy, highlighting the need to evaluate no-CoT training strategies in the context of the full post-training pipeline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑