QVAC Genesis III:面向高效语言模型预训练的大规模、高质量开放合成STEM语料库
QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
- Tether Data, S.A. de C.V. d.b.a. Tether AI Research(Tether Data, S.A. de C.V.(以Tether AI Research名义运营))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出1914.3亿token的STEM合成语料库QVAC Genesis III,通过双生成策略和LLM-as-parser评估,在1.7B模型上显著优于现有开源数据,提升ARC和MMLU STEM性能。
AI中文摘要:
高质量预训练数据是面向边缘AI和端侧部署的教育及STEM专用语言模型的关键瓶颈,这些场景中token预算受到严格限制。尽管主要机构在私有语料库上训练越来越大的模型,但开放生态缺乏针对STEM的合成数据集,能够为小模型高效提供高单token学习价值。为弥补这一空白,我们提出QVAC Genesis III,一个包含1914.3亿token、覆盖19个领域、多个难度级别和不同教育风格的STEM多领域合成语料库。QVAC Genesis III通过双生成策略构建,该策略使用弱边缘规模学生模型作为信号进行针对性教师蒸馏:学生模型的失败被转化为纠正性解释,而其成功则被扩展为对所有答案选项的对比性选项级推理。我们进一步引入一种LLM-as-a-parser评估协议,从自由格式输出中提取最终答案,并同时跟踪准确率和答案有效性。为验证QVAC Genesis III数据的有效性,我们使用17亿参数模型进行受控从头消融实验,结果表明,使用QVAC Genesis III训练的模型在ARC、GPQA Diamond和MMLU STEM基准上始终优于使用开源合成语料库Cosmopedia-v2训练的模型以及公开的Cosmo-1B模型,在ARC-E上最高提升+28.57%,在ARC-C上最高提升+21.35%,同时有效答案率最高达到99.45%。
英文摘要:
High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the open ecosystem lacks STEM-focused synthetic datasets that deliver high per-token learning value efficiently for small models. To address this gap, we introduce QVAC Genesis III, a 191.43B-token, STEM-focused multi-domain synthetic corpus covering 19 domains across several difficulty levels and different educational styles. QVAC Genesis III is built via a dual generation strategy that performs targeted teacher distillation using a weak edge-scale student model as signal: the student's failures are converted into corrective explanations, while its successes are expanded into contrastive option-level reasoning over all answer choices. We further introduce an LLM-as-a-parser evaluation protocol that extracts final answers from free-form outputs and tracks both accuracy and answer validity. To validate the effectiveness of our QVAC Genesis III data, we conduct controlled from-scratch ablations with 1.7B-parameter models, showing that models trained with QVAC Genesis III consistently outperform both models trained with the open-source synthetic corpus Cosmopedia-v2 and the publicly released Cosmo-1B model across ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% on ARC-E and +21.35% on ARC-C, while reaching a Valid Answer Rate of up to 99.45%.