发表机构
Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何从专家行动中恢复思维链,提出LeAct方法,通过优化潜在变量,让学生为专家行动采样候选思维链并保留有效链。该方法在多个博弈和机器人基准测试中表现出色,使专家系统成为基础模型推理教师的新来源。
AI 中文摘要
现代推理模型依赖于推理数据,目前这些数据来自人工标注或从更强的语言模型中提炼。然而,丰富且未被充分利用的监督来源在于专家系统,其能在不同领域常规地产生近乎最优的行动,但这些专家不会记录行动背后的思维链。我们将恢复思维链视为一个潜在变量,并研究如何仅从行动中恢复它。我们的方法LeAct通过优化潜在变量,让学生为每个专家行动采样候选思维链,保留能提高恢复行动概率的思维链。在多个规模的不完全信息博弈和一个模拟机器人基准测试中,LeAct在小型可枚举博弈上达到求解器的数值下限;在更大规模上,它比最强的专家迭代基线接近求解器5倍;在翻牌德州扑克中,LeAct以每手牌多赢60个大盲注的优势直接获胜,在机器人测试中,它是唯一比直接模仿有改进的训练方法。我们提出了一个有原则的框架及结果:专家系统成为基础模型全新的推理教师来源。
英文摘要
Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem provers), which routinely produce near-optimal actions across diverse domains. But these experts are silent: they commit to an action without writing down the chain of thought (CoT) behind it. Recovering that CoT as natural-language reasoning would distill expert knowledge into a student that generalizes beyond the demonstrated actions. We treat it as a latent variable and study how to recover it from the action alone. Our approach, LeAct (Learning to reason from Actions), optimizes this latent variable: the student samples candidate CoTs for each expert action, and we retain those that measurably improve its own probability of recovering the action. Across imperfect-information games at multiple scales and a simulated robotics benchmark, LeAct reaches the solver's numerical floor on small enumerable games. At larger scale, it is $5\times$ closer to the solver than the strongest expert-iteration baseline. At Flop Hold'em ($\sim 10^9$ infosets), LeAct wins head-to-head by $+60$ mbb/g, and on the robotics probe it is the only training recipe that improves on direct imitation. We present a principled framework and the result: expert systems become a categorically new source of reasoning teachers for foundation models.
Comments27 pages, 3 figures, 11 tables