arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09425cs.CLcs.AI

Edu-QuRating:基于蒸馏成对判断的多维教育数据策展

Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

Oliver G. B. Garrod, Robin A. A. Ince, Meng Liu, Mohamed Huti, Moritz Boos, Amy Waldock, Dominic Andrews, Romana Alonso-Kropil, Paul Atherton

首次发表
浏览论文内容

中文总结 AI 辅助

Edu-QuRating提出多维教育数据评分流程,通过蒸馏成对偏好训练评分器,提升小模型预训练和GRPO后训练效果,优于FineWeb-Edu基线。

中文摘要 AI 辅助

教育数据过滤器已成为提升语言模型预训练的一种实用方法,但大多数过滤器将教育价值视为单一的标量属性。这对于某些应用而言可能过于宽泛,尤其是当数据集本身已具有高密度的教育材料时。有用的学习材料需要准确、引人入胜、结构良好,并适合目标受众和应用场景(例如面向学习者与面向教师)。遵循QuRating(Wettig等人,2024年),我们引入了Edu-QuRating:一个用于多维教育数据评分和策展的流程。Edu-QuRating定义了特定于教育的评分标准,使用LLM裁判对采样的文档对进行标注,并将这些成对偏好蒸馏为可复用的Edu-QuRaters,后者能够根据一组教育标准对单个文本块进行评分。在两种序列分类基础模型和六项教育标准中,最佳的Edu-QuRater以平均准确率0.917恢复了留出的GPT-4.1-mini成对判断。随后,我们将所得的评分器应用于两个应用场景。首先,我们研究了Edu-QuRaters在语料库过滤中用于改进小型语言模型预训练的潜力。我们对322.25M篇FineWeb-Edu-Fortified文档进行了评分,以获得过滤后的预训练混合数据。在匹配的单次运行预训练比较中,使用基于Edu-QuRating的混合数据训练的模型在九个基准上的总体聚合准确率高于FineWeb-Edu基线,且增益集中在特定任务上。其次,我们将Edu-QuRater分数用作GRPO后训练中的奖励项。在留出的成对裁判评估中,将Edu-QuRater与答案结构奖励相结合产生的响应在教学质量和对指令遵循两方面均优于Qwen3-4B基础模型。

英文摘要

Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of educational material. Useful learning material needs to be accurate, engaging, well structured, and appropriate for the intended audience and application (e.g. learner- vs teacher-facing). Following QuRating (Wettig et al. 2024), we introduce Edu-QuRating: a pipeline for multi-dimensional educational data scoring and curation. Edu-QuRating defines education-specific rubrics, uses an LLM judge to label sampled document pairs and distills those pairwise preferences into reusable Edu-QuRaters, which can score individual text chunks on a set of educational criteria. Across two sequence-classification base models and six educational criteria, the best Edu-QuRater recovers held-out GPT-4.1-mini pairwise judgements with mean accuracy 0.917. We then apply the resulting scorers in two applications. First, we investigate the potential of Edu-QuRaters for corpus filtering to improve pretraining of small language models. We scored 322.25M FineWeb-Edu-Fortified documents to obtain a filtered pre-training mixture. In matched single-run pre-training comparisons, models trained with Edu-QuRating-based mixtures reached higher observed aggregate accuracy across nine benchmarks than the FineWeb-Edu baseline, with gains concentrated in particular tasks. Second, we used Edu-QuRater scores as reward terms for GRPO post-training. In held-out pairwise judge evaluations, combining Edu-QuRater and answer-structure rewards produced responses preferred to the Qwen3-4B base model on both pedagogical quality and instruction following.

发表机构

  • Fab AI

机构由 AI 辅助整理,请以论文原文为准。

↑