AI 中文总结
EvalConvoLearn是评估辅导对话中学习者模拟的开源框架,从学习行为与对话质量两个维度开展评估,验证了其在辅导对话数据集上的有效性并公开了代码。
AI 中文摘要
对话式学习者模拟是用于测试学习理论、评估教学材料和自动导师或驱动可教学智能体的宝贵工具。近年来,大语言模型(LLM)实现了与模拟学习者更丰富、更自然的交互;然而,目前尚无开源框架用于评估此类模拟是否忠实地再现真实学习者行为。我们推出EvalConvoLearn,一个从两个维度评估学习者模拟的开源框架:学习行为(技能条件掌握结果)和对话质量(言语行为、错误类型分布、提问率、轮次长度)。EvalConvoLearn通过将指标基于真实辅导对话数据集,并将生成的导师响应锚定到现有导师话语中,来衡量模拟学习者与数据中观察到的答案分布的接近程度。该框架在一个辅导对话数据集上得到验证,包括两个基于LLM的学习者模拟的结果,以及已发布的GitHub代码。
英文摘要
Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tutoring dialogues, including results for two LLM-based learner simulations, and the published GitHub code.
CommentsPoster at the Impactful and Responsible AI Systems for Education workshop, as part of the Festival of Learning 2026