arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StudentSim:基于大语言模型的学生模拟器训练框架

StudentSim: Training LLM-based Student Simulators

Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai, Jianfeng Gao

arXiv 2609.01591首次发表:更新:

发表机构

Microsoft Research; University of Illinois Urbana-Champaign(微软研究院; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出 StudentSim 框架,结合标准化评估协议 StudentSimEval,在国际象棋、写作、数学领域训练出优于 GPT-5.4 的学生模拟器,用作奖励模型可生成更优质的个性化辅导系统。

AI 中文摘要

AI 辅导系统的最大价值在于能适配每个学生的优势、劣势和偏好的指导方式,但关于哪种指导对哪个学生有效的证据,从真实学习者那里收集既稀少、缓慢又成本高昂。学生模拟器可作为代理提供这种信号,但现有方法存在局限:状态追踪模型能拟合学生行为,但难以处理解释或修正内容;而基于大语言模型(LLM)的角色扮演虽能流畅遵循指导,却无法可靠匹配被模仿学生的能力。我们提出 StudentSim,这是一种训练框架,它通过汇集训练后再按学生个性化调整的方式,将稀疏的学生个体数据转化为个性化模拟器。生成的模拟器既能镜像学生自身的反应,又能在辅导指导下更新自身。我们还推出 StudentSimEval,这是一个标准化评估协议,涵盖国际象棋、英语第二语言写作和数学三个领域的60名学生,使用的是供研究共享的去标识化公开学习者数据集。StudentSimEval 衡量两个指标:行为保真度(F,即模拟器匹配学生反应的程度)和指导响应性(R,即模拟器在辅导指导下更新的难易程度),所有方法均在相同记录上拟合和评估。在所有三个领域,StudentSim 在两个指标上均优于 GPT-5.4:在国际象棋领域,StudentSim 的 F 值达0.51、R 值达0.91,而 GPT-5.4 分别为0.23和0.72,Maia2 则分别为0.45和0.27。作为概念验证,将 StudentSim 用作辅导强化学习的奖励模型,生成的国际象棋辅导系统经人类专家评估,比无强化学习基线和基于 GPT-5.4 模拟器奖励训练的辅导系统更准确、指导更得当、个性化程度更高。代码可在该 https URL 获取。

英文摘要

AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑