发表机构
Swiss Federal Institute of Technology in Lausanne (EPFL)(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EduFair-Bench通过多领域题库和受控模拟,系统评估LLM导师在不同学生人口统计特征下的教学公平性,发现模型能力与公平性正交,且教学RL训练仅重新分配而非消除偏差。
AI 中文摘要
大型语言模型(LLM)越来越多地被部署为导师,但尚不清楚它们是否对所有学生都能提供同等质量的辅导。我们引入了\ extbf{EduFair-Bench},一个用于审计LLM导师教学公平性的基准——即辅导质量是否随学生人口统计特征而系统性变化。EduFair-Bench将多领域题库(数学、物理、化学)与受控模拟相结合,在该模拟中,一个固定的LLM学生与每位导师在跨越四个维度的九个不同人口统计水平上进行交互,这四个维度包括:性别、移民背景、第一语言和社会经济地位(SES)。辅导质量通过五个回合级教学指标和四个对话级维度进行评分,使用一个经过180个导师回合的三标注者共识验证的LLM评判器。偏差通过配对Wilcoxon符号秩检验和自助法效应量置信区间来衡量。两个消融实验(通过名字传达的人口统计线索;导师与学生之间冲突的人口统计信息)将导师驱动的偏差与学生驱动的偏差区分开来。在五位导师中,我们发现模型能力和人口统计公平性在很大程度上是正交的:最小的模型最为一致,而四位能力更强的导师都表现出广泛的人口统计差距,且没有明确的能力与公平性排序;针对教学法的RL训练重新分配而非消除了偏差;与语言和移民相关的线索产生的差距大于与性别和社会经济地位相关的线索。
英文摘要
Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.