arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DomainPilot:用于高效语言模型微调的域级损失引导两阶段数据混合优化

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

He Zhang

arXiv 2607.22769首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型训练数据问题,提出DomainPilot框架,通过域级损失引导两阶段数据混合优化,经令牌级监测、缩放与混合定律引导优化,在Qwen3-1.7B模型微调中验证,有效提升指标且不增成本。

AI 中文摘要

大语言模型的训练效率受训练数据质量和组成的根本限制。现有动态数据调度方法在工业规模预训练和监督微调中存在关键局限。我们提出DomainPilot,一个域级损失引导两阶段数据混合优化框架。它引入令牌级域损失监测,基于此提出缩放定律引导的粗优化阶段和混合定律引导的细优化阶段,通过基于补丁的架构实现,在Qwen3-1.7B模型的监督微调中验证,优化混合在多项指标上有提升且不增加数据量和训练成本。

英文摘要

The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining and supervised fine-tuning (SFT): data selection incurs prohibitive O(N) costs on terabyte-scale corpora, mixture optimization schemes introduce severe I/O bottlenecks or require training auxiliary reference models, and sample-level reweighting strategies rely on loss signals that conflate noise, difficulty, and novelty. We present DomainPilot, a domain-level loss-guided two-stage data mixture optimization framework. DomainPilot introduces token-level domain loss monitoring to capture per-domain learning dynamics during training without halting the data pipeline. Building on these signals, we propose a Scaling Law guided coarse optimization stage that fits domain-specific convergence curves and derives a principled prior for mixture adjustment. A subsequent Mixing Law guided fine optimization stage refines the mixture by modeling cross-domain interaction effects through controlled sweep experiments. The entire mechanism is realized via a patch-based architecture that injects domain-aware loss computation into existing training frameworks (e.g., MindSpeed/Megatron-LM) with only ~30 lines of framework-specific adapter code. We validate DomainPilot on the Qwen3-1.7B model during SFT. Compared to the original data mixture, our optimized mixture achieves improvements of +2% on MMLU-Redux, +1.8% on AIME24, +3.8% on LiveCodeBench v5, and +3.6% on BFCL v3, without increasing total data volume or training cost. These results demonstrate that domain-level training signals provide an effective, lightweight alternative to expensive data selection or auxiliary model training for mixture optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑