AI 中文总结
本研究探究RL与SFT微调模型数学推理性能差异的机制,发现RL模型的表征更具线性可分性、深层重要性更高,其token分配变异性或反映在线策略推理的分布。
AI 中文摘要
越来越多研究表明,通过强化学习(RL)训练的大型推理模型在数学推理任务上的表现优于经监督微调(SFT)的对应模型,但这种优势的机制基础仍不明确。因此,本研究提出核心问题:何种内部表征差异使RL模型具备更优性能?研究提供了两条相互印证的证据:其一,针对分层隐藏状态训练的线性探测显示,RL模型在预测答案正确性时的准确率往往高于SFT模型,表明其表征更具线性可分性与结构性;其二,均值消融研究表明,RL模型形成了层级架构,深层的重要性逐步提升,而SFT模型则将重要性均匀分布在各层。这些发现共同证明,RL训练从根本上重构了模型表征与处理推理问题的方式。最后,本研究分析了不同问题下重复采样的token计数变异性以评估自适应计算分配,虽观察到部分RL微调模型的变异性高于SFT对应模型,但其他RL模型则表现出强一致性,说明token分配可能更多取决于整体训练流程,而非仅由RL与SFT的差异决定。研究认为,这种token分配变异性反映了合理的在线策略推理的分布,可用于区分哪些模型具备稳定策略,哪些则存在欠确定、潜在不可识别的解行为。
英文摘要
Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models' superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.
CommentsSecond Workshop on XAI4Science, AAAI 2026