StudentBench:AI与人类辅导在GRE学习收益上效果相当
StudentBench: AI and human tutoring yield equivalent GRE learning gains
浏览论文内容
中文总结 AI 辅助
本研究提出StudentBench平台,通过大规模实验证明AI辅导在GRE学习收益上与人类辅导等效,且成本极低,并揭示了AI回复速度与学习成效的正相关关系。
中文摘要 AI 辅助
人工智能为增强人类能力提供了前所未有的机遇,然而前沿进展主要聚焦于提升模型能力。我们推出了StudentBench,一套AI教学评估工具及一个公共平台,该平台支持大规模数据收集,包含超过175,000条学生与AI的交互消息,用于研究大语言模型(LLMs)是否能够产生与人类辅导相当的学习收益。利用StudentBench,我们在2,383名接受AI辅导、人类辅导或未接受辅导的人类参与者中,测量了他们在定量和语文GRE题目上的学习收益。我们证实,在GRE学习收益方面,AI辅导在统计上与专家人类辅导等效(p = .015),并且在七个GRE领域中,表现最佳的AI辅导者在其中五个领域的平均表现超过了人类辅导者。在第二项研究中,专家人类辅导者通过2,028次成对评分评估,比较了LLM生成的教案和练习题。两项研究共同从以下五个方面清晰地区分了AI辅导者:(1)教案规划,(2)练习题创作,(3)对话式教学法,(4)成本,以及(5)参与度。令人惊讶的是,一位AI辅导者实现了与人类辅导等效的学习收益(p = .044),而成本却低918倍(AI每提升一个百分点花费0.0052美元,人类则需4.81美元)。在定量GRE会话中,更快的AI回复与学生发送更多消息相关,更多消息与更多正确练习相关,而更多正确练习与更大的学习收益相关(所有p值均小于.002)。StudentBench平台可免费访问,网址为https://studentbench.org。
英文摘要
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). To support future research, we open-source the de-identified data collected in our studies.