arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ELBench:面向教育场景的大语言模型多维基准

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Guangtao Zhai, Aimin Zhou

arXiv 2608.09548首次发表:更新:

发表机构

East China Normal University; The Hong Kong University of Science and Technology (Guangzhou); Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory(华东师范大学; 香港科技大学(广州); 上海交通大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ELBench是首个评估教育大模型四项核心要求的综合基准,评估发现通用模型综合表现相近但模块优势不同,中国模型在安全性模块领先,教育专用模型未在教育模块占优,高阶培养存在系统性盲点。

AI 中文摘要

大语言模型正越来越多地被部署在教育领域,担任导师、教学助手和内容生成器等角色。这些角色对模型提出了普通问答任务所没有的要求:可用的面向教育场景的模型需同时具备准确性、敏感提示下的安全性、教学实用性,且符合教学目标。现有基准大多单独评估这些要求,因此没有任何一个基准能将面向教育场景的适用性作为综合维度进行评估。我们推出ELBench,这是首个在同一协议下评估所有四项要求(通用能力、安全性与可信性、基础教育、高阶培养)的基准,结合了精选的公开资源与新合成的安全和培养数据。我们评估了九种模型,其中七种是前沿通用系统,两种是教育专用变体,并报告了三项发现:第一,模块级别的特征比单一的综合分数更具信息量;排名前六的模型在综合分数上无统计学差异,但它们的模块优势差异显著,且安全性与实际教学呈负相关(r = -0.83)。第二,中国开发的模型在安全性模块中表现领先,该模块是套件中最具区分度的部分;这一优势在特定地区的规范内容上最为明显,在通用危害内容上虽有所缩小但并未消失。第三,两种教育专用模型在两个教育模块中均未领先,且在高阶培养任务上,所有模型都存在系统性盲点:在结构化判断任务中,它们都倾向于非参考选项,更偏好教学风格而非符合既定目标,因此该模块得分普遍较低,无法区分模型。这引发了一个问题:领域后训练是否能在教育任务上跟上前沿系统的步伐,但并未解决该问题。

英文摘要

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.

Comments13 pages, 6 figures, 8 tables. Benchmark data: https://huggingface.co/datasets/ZeroLoss-Lab/ELBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑