语言模型在强化学习后训练下的组合推理
Compositional Reasoning in Language Models under Reinforcement Learning Post-Training
浏览论文内容
中文总结 AI 辅助
本文提出依赖图框架形式化组合推理,通过数据结构任务发现分解到组合训练迁移的不对称性,并提供理论解释及现实工具调用基准的初步验证。
中文摘要 AI 辅助
组合推理对于现实世界的问题解决至关重要:由于训练数据必然有限,模型必须通过以新的方式组合已学技能来进行泛化。尽管诸如强化学习(RL)之类的后训练方法已显著提升了语言模型(LMs)的推理能力,但它们对组合推理的影响仍不太清楚。我们提出了一个依赖图框架来形式化组合推理,产生了三个复杂度递增的组合层级。在实证中,我们用数据结构任务实例化该框架,这些任务提供了确定性的奖励计算和清晰的组合结构。我们发现了一致的分解到组合的不对称性:分解技能训练不能可靠地迁移到组合任务,而组合任务训练则更容易迁移回分解任务。我们为这种不对称性提供了理论解释,并进一步评估了在长度外推、结构分布偏移以及迁移到需要未见技能的任务下的组合泛化。最后,我们展示了一项关于现实世界工具调用基准的初步研究,表明分解到组合的不对称性可以扩展到实际场景。
英文摘要
Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.
发表机构
- Stanford University(斯坦福大学)
- Amazon AGI Labs(亚马逊AGI实验室)
机构由 AI 辅助整理,请以论文原文为准。