学习教练以进行体验式学习
Learning to Coach for Experiential Learning
浏览论文内容
中文总结 AI 辅助
提出L2C框架,训练LLM作为教练从演员轨迹中提取经验知识,通过同实例和跨实例奖励优化,在数学推理和文本游戏中优于自我改进,并有效利用推理计算。
中文摘要 AI 辅助
语言模型可以从经验中学习,但原始的解决方案轨迹往往过长且充满噪声,难以提供有效的指导。在这项工作中,我们提出了学习教练(L2C)框架,该框架训练一个专门的LLM-as-a-Coach(大语言模型作为教练),从演员模型的先前轨迹中提取可操作的体验知识。演员模型保持冻结,而LLM-as-a-Coach被训练以最大化由演员模型在指导下回答的正确性所给出的奖励。我们研究了两种这样的奖励:同实例奖励,它改进对原始问题的后续回答;以及跨实例奖励,它引出可迁移到其他实例的知识。在数学推理和交互式文本游戏任务中,L2C始终优于自我改进和未经训练的LLM-as-a-Coach。运行更多迭代的体验式学习进一步提高了准确性,并且比扩大演员模型的解码预算更有效地利用额外的推理计算。训练好的LLM-as-a-Coach还能迁移到分布外任务,并根据其指导的具体演员模型调整指导方式。
英文摘要
Language models can learn from experience, but raw solution trajectories are often too long and noisy to provide effective guidance. In this work, we propose Learning to Coach (L2C), a framework that trains a dedicated LLM-as-a-Coach to extract actionable experiential knowledge from an actor model's previous trajectory. The actor remains frozen, while the LLM-as-a-Coach is trained to maximize a reward given by the correctness of the actor's guided response. We study two such rewards: a same-instance reward, which improves subsequent responses on the original problem, and a cross-instance reward, which elicits knowledge that transfers to other instances. Across mathematical reasoning and interactive text-games, L2C consistently outperforms self-refinement and an untrained LLM-as-a-Coach. Running experiential learning for more iterations further improves accuracy and uses additional inference compute more effectively than enlarging the actor's decoding budget. The trained LLM-as-a-Coach also transfers to out-of-distribution tasks and adapts its guidance to the specific actor it coaches.
发表机构
- Microsoft Research(微软研究院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。