arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据混合作为混合实验:大语言模型预训练的响应面方法与最优设计

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

Yicheng Mao, Hongru Du

arXiv 2608.23922首次发表:更新:

发表机构

University of Calgary; University of Virginia(卡尔加里大学; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将LLM预训练的数据混合视为混合实验,采用Scheffé响应面模型构建模型鲁棒的$\boldsymbol{\textit{I}}$-最优设计,可解释混合响应并提升代理实验效率,验证了其统计优势。

AI 中文摘要

数据混合是大语言模型(LLM)预训练中的核心设计问题:在给定固定token预算的情况下,从业者必须决定为每个领域分配多少数据。近期基于代理的方法通过在候选混合数据上训练小型模型、拟合响应模型,并利用该响应结果为更大规模的训练选择混合数据。我们表明,该工作流程具有经典混合实验的结构:在此视角下,数据领域是混合组分,token占比是组分比例,代理训练运行是实验设计点,验证损失则定义了概率单纯形上的响应面。我们采用稀疏二阶Scheffé响应面模型开发该公式,并为代理数据混合实验构建了模型鲁棒的$\boldsymbol{\textit{I}}$-最优设计。以RegMix作为实证案例研究,我们展示了该框架如何既解释观测到的混合响应,又设计更高效的代理实验。Scheffé分析表明,领域价值具有强关联性:多个在加性效应下表现较弱的领域,可通过成对交互作用变得有利,尤其是与网络衍生文本的组合。稀疏Scheffé模型保留了跨模型尺度的混合排序,且在提供加性和交互效应显式分解的同时,仍与灵活的机器学习预测器具有竞争力。在一项基于观测代理训练响应校准的模拟研究中,模型鲁棒的$\boldsymbol{\textit{I}}$-最优设计在移除约25%的原始代理运行后,仍能恢复相关的混合排序。这些结果表明,LLM数据混合不仅应被视为一个预测问题,还应被视为一个实验设计问题,其中代理混合数据本身可被选择以提高统计效率。

英文摘要

Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by training small models on candidate mixtures, fitting a response model, and using the response to select mixtures for larger-scale training. We show that this workflow has the structure of a classical mixture experiment. Under this view, data domains are mixture components, token shares are component proportions, proxy-training runs are experimental design points, and validation loss defines a response surface over the probability simplex. We develop this formulation using sparse second-order Scheffé response-surface models and construct model-robust $\mathcal{I}$-optimal designs for proxy data-mixing experiments. Using RegMix as an empirical case study, we demonstrate how the framework can both interpret observed mixture responses and design more efficient proxy experiments. The Scheffé analysis shows that domain value is strongly relational: several domains that are weak under additive effects become favourable through pairwise interactions, especially through combinations with web-derived text. The sparse Scheffé model preserves mixture rankings across model scales and remains competitive with a flexible machine-learning predictor while providing an explicit decomposition of additive and interaction effects. In a simulation study calibrated to observed proxy-training responses, model-robust $\mathcal{I}$-optimal designs recover the relevant mixture ordering after removing about 25\% of the original proxy runs. These results suggest that LLM data mixing should be treated not only as a prediction problem, but also as an experimental-design problem in which the proxy mixtures themselves can be chosen to improve statistical efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑