arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.16340cs.CLcs.AI

思考中的思考:评估后训练语言模型的推理能力

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

  • Indian Institute of Technology Roorkee(印度理工学院罗奥克学院)
  • Boston University(波士顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Pratham Singla, Shivank Garg, Ayush Singh, Ishan Garg, Ketan Suhaas Saichandran

更新

AI总结:

研究评估后训练语言模型的推理能力,通过三个核心能力评估其意识、泛化和推理与输出的对齐情况,发现强化学习模型在意识和泛化上优于SFT模型,但推理与输出对齐较差。

AI中文摘要:

近期后训练技术的进步使大型语言模型(LLMs)能够通过生成补充规划令牌来应对复杂逻辑任务。这引发了核心问题:这些模型是否意识到自己所学和所思?我们定义了三项核心能力:(1)对学习潜在策略的意识,(2)这些策略在不同领域中的泛化能力,以及(3)内部推理轨迹与最终输出的一致性。我们通过多个任务对这些能力进行了实证评估,每个任务均要求学习不同的策略。此外,我们对比了通过监督微调(SFT)、直接策略优化(DPO)和组相对策略优化(GRPO)后训练的模型。研究发现,强化学习训练的模型不仅在意识和对新结构相似任务的泛化能力上优于SFT模型,而且在推理轨迹与最终输出的一致性上表现较差,尤其在GRPO训练的模型中表现最为显著。

英文摘要:

Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensive tasks through the generation of supplementary planning tokens. This development raises a fundamental question: Are these models aware of what they "learn" and "think"? To address this, we define three core competencies: (1) awareness of learned latent policies, (2) generalization of these policies across domains, and (3) alignment between internal reasoning traces and final outputs. We empirically evaluate these abilities on several tasks, each designed to require learning a distinct policy. Furthermore, we contrast the profiles of models post-trained via Supervised Fine-Tuning (SFT), Direct Policy Optimization (DPO), and Group Relative Policy Optimization (GRPO). Our findings indicate that RL-trained models not only demonstrate greater awareness of their learned behaviors and stronger generalizability to novel, structurally similar tasks than SFT models but also often exhibit weak alignment between their reasoning traces and final outputs, an effect most pronounced in GRPO-trained models.

↑