arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06446cs.CLcs.AI

教学困境:通过多轮强化学习学会教学

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

Jakub Macina, Manu Kapur, Mrinmaya Sachan

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM辅导者易直接告知答案的问题,提出掩蔽近迁移后测与二元奖励门,开发多轮RL方法Eduardo,训练出高效且教学能力强的模型。

中文摘要 AI 辅助

经过训练用于回答问题的的大型语言模型(LLMs)天生不擅长教学。针对模拟学生进行强化学习(RL)是提升其教学能力的有效途径,但现有经RL训练的辅导者以辅导问题上的学生成功作为奖励,且辅导者的言语仍保留在上下文中。因此,最易提升奖励的方式是直接告诉学生答案,这需要调整惩罚以减少直接告知。借鉴学习科学,我们引入一种掩蔽近迁移后测:学生在辅导者话语被掩蔽的情况下,接受辅导问题的一个未见变体测试,因此奖励只能通过学生在其自身回合中书写的内容来提升。这抑制了学生的认知卸载,并允许将连续惩罚替换为两个二元奖励门(辅导者回答的事实正确性、无解决方案移交)。留一消融实验表明,单独的学习增益奖励无法区分教学与直接告知:奖励门减少了解决方案移交,而近迁移后测提升了域外迁移能力。利用这些奖励设计,我们开发了Eduardo,一种用于训练LLM辅导者的多轮RL方法,并使用它从两种不同的LLM架构训练了4B、9B、14B和27B模型。我们后训练的Eduardo-27B模型在MathTutorBench上匹配Gemini-3.1-Pro,在TutorMoments上匹配Claude Opus 4.8,且思考令牌数比前沿模型少2.4-6.2倍,这对交互式辅导至关重要。尽管奖励中未明确提及,模型使用推动论证教学动作的频率增加了一倍以上,而支持消退(例如分配独立工作)——其收益超出单问题对话片段——则被训练掉了。我们开源了训练环境、包含8,671个问题的近迁移数据集以及训练好的模型,以供进一步开发。

英文摘要

Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.

发表机构

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑