Sherpa:教大语言模型自适应教学
Sherpa: Teaching LLMs to Teach Adaptively
浏览论文内容
中文总结 AI 辅助
提出多轮强化学习框架Sherpa,通过模拟多样学生原型并直接最大化学习成果,训练LLM教师自适应教学,平均提升学生表现20.5个百分点,教学得分从52.5%升至79.2%。
中文摘要 AI 辅助
大语言模型(LLMs)已成为越来越强大的问题解决者,但能够解决问题并不等同于能够教授问题。现有的将LLMs训练为教师的方法依赖于演示、偏好数据或预定义的教学标准,这些标准规定了良好教学应是什么样的。然而,这些信号往往并未基于个体学生的学习成果,而有效的教学策略在不同学习者之间可能存在显著差异。为解决这一问题,我们提出了Sherpa,一个多轮强化学习框架,该框架利用LLMs实例化多个基于不同学习偏好条件化的学生原型,并训练教师模型通过直接最大化其学习成果来调整其教学。使用Sherpa训练的教师LLMs在所有原型上平均将受教学生的表现提升了20.5个百分点。在MathTutorBench的评估下,Sherpa将整体教学得分从52.5%提升至79.2%,表明教学回应更优。我们的人类研究显示,在79.6%的成对比较中,训练后的教师模型优于基础模型。综上,Sherpa训练LLM教师适应多样化的模拟学生,并使其与人类教师更好地对齐,为AI导师教授真实学生铺平了道路。
英文摘要
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
发表机构
- Stanford University(斯坦福大学)
- Georgia Tech(佐治亚理工学院)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。