EduClaw-Bench:面向教学大型语言模型智能体的长周期基准测试集(含模拟学习者)
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
浏览论文内容
中文总结 AI 辅助
研究人员推出EduClaw-Bench长周期基准测试集,结合模拟学习者评估LLM智能体辅导效果,发现辅导质量取决于基础模型与智能体适配器的结合,且多数组合无法维持全程良好辅导。
中文摘要 AI 辅助
大型语言模型(LLM)为从辅导到作文评分的各类教育应用提供支持,但每种应用仅针对单一任务,且仅在近期这些单点解决方案才被整合到在学习管理系统(LMS)上运行的智能体中。然而辅导是长周期任务,因为学习者的进步发生在数天乃至数周的时间里,而非单轮交互,且目前尚无基准测试集能对智能体辅导的持续关系进行评估。我们推出 EduClaw-Bench,这一基准测试集将智能体辅导者置于与基于知识追踪(KT)的模拟学习者的连续30天关系中,该模拟学习者的知识概念掌握度来自于用真实学生数据训练的KT模型,其答案由该掌握度驱动,并在55种场景中被探测学习增益。每个智能体在三个主要维度(学习增益、响应性和有用性)以及两个课程设计维度(加涅和罗森夏因原则)上被评分,其中有用性和课程维度由三名跨家族的LLM评判员组成的小组评判。对10种智能体适配器在三个基础模型层级上进行评估,得出了两个单层级、单会话评估无法得出的发现:第一,辅导质量属于基础模型和智能体利用方式的结合,而非仅属于其中一方;第二,几乎没有任何组合能在整个周期内维持良好的辅导效果。校准检查(ECE=0.049)和现场课堂研究证实,模拟学习者及其测量结果与现实相符。我们的工作是迈向面向未来教育的可信AI辅导者的一步。
英文摘要
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS). Yet tutoring is long-horizon, since a learner improves over days and weeks rather than in a single turn, and no benchmark evaluates an agent tutor across a sustained relationship. We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagné and Rosenshine), with helpfulness and the curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters over three base-model tiers yields two findings that single-tier, single-session evaluation cannot reach. First, tutoring quality belongs to the base model and the agent harness together rather than either alone. Second, almost no combination sustains good tutoring over the full horizon. A calibration check ($\text{ECE}=0.049$) and a live-classroom field study confirm that the simulated learner and its measurements track reality. Our work is a step toward trustworthy AI tutors for future education.
发表机构
- Korea University Sejong Campus(高丽大学世宗校区)
- Opentutorials
- Indiana University(印第安纳大学)
- Gyeonggi Institute of Education(京畿教育大学)
- Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。