arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23088cs.CL

OmniEdu:面向学习与教学的开源基础模型

OmniEdu: Open Foundation Models for Learning and Teaching

Hao Liang, Qihan Lin, Meiyi Qiang, Linzhuang Sun, Hengyi Feng, Mingrui Chen, Sizhe Qiu, Wentao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

OmniEdu提出面向K-12教育的能力均衡指令微调方法,整合百种资源构建69,999个示例,微调4B至27B模型,在多项教育基准上显著提升,验证了能力导向数据对教育任务适配的有效性。

中文摘要 AI 辅助

教育基础模型必须能够解决问题、理解课程结构、诊断学习者困难,并提供适当的教学支持。现有的教育语言模型通常侧重于问题求解或辅导中的某一方面,其训练数据混合按来源或任务组织,而非按能力组织。我们提出OmniEdu,一个面向K-12学习和教学的开源基础模型家族。其指令微调语料库整合了超过100种教育资源与通用指令来源,围绕四种能力组织:学科能力、课程基础、诊断推理以及教学行动与脚手架支持。我们的流程整合了确定性清洗、语义审计与重写、任务特定质量评分、基于token预算的多样性选择以及教学指令分配。该流程生成了69,999个示例和15.96M个监督响应token,其中包括60,951个教育特定示例。我们微调了4B、9B和27B模型,并评估了课程基础、K-12问题求解、教学辅导以及通用能力。面向教育的微调一致地提升了所有三种教育基准组在多个模型规模上的表现。OmniEdu-27B在K12-Bench上达到63.12%的EM和76.69%的F1,在MathFish上达到85.89%,在EDUMATH上达到86.95%,在MathTutorBench的Scaffold设置中达到78.74%。它还在LongTutor上取得了评估模型中最高的Teaching平均值,为3.02。这些结果证明了经过精心策划的、能力均衡的监督数据对于将通用语言模型适配到涵盖问题求解、课程理解和教学支持的教育任务中的价值。

英文摘要

Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capability. We present OmniEdu, an open family of foundation models for K-12 learning and teaching. Its instruction-tuning corpus combines over 100 educational resources and general instruction sources, organized around four capabilities: subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action and scaffolding. Our pipeline integrates deterministic cleaning, semantic auditing and rewriting, task-specific quality scoring, token-budgeted diversity selection, and pedagogical instruction assignment. It yields 69,999 examples and 15.96M supervised response tokens, including 60,951 education-specific examples. We fine-tune 4B, 9B, and 27B models and evaluate curriculum grounding, K-12 problem solving, and pedagogical tutoring, alongside general capability. Education-oriented tuning consistently improves all three educational benchmark groups across model scales. OmniEdu-27B achieves 63.12% EM and 76.69% F1 on K12-Bench, 85.89% on MathFish, 86.95% on EDUMATH, and 78.74% in MathTutorBench's Scaffold setting. It also achieves the highest Teaching average on LongTutor among the evaluated models, at 3.02. These results demonstrate the value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support.

发表机构

  • Peking University(北京大学)
  • University of the Chinese Academy of Sciences(中国科学院大学)
  • Zhongguancun Academy(中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

↑