arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38011cs.LGstat.ML

数据混合何时能改进缩放律?来自高维回归的洞见

When do data mixtures improve scaling laws? Insights from high-dimensional regression

Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过高维回归理论分析,揭示了数据混合改进缩放律的条件:需谱与相对样本量特定匹配,并在语言模型实验中验证了该现象。

中文摘要 AI 辅助

现代机器学习系统在来自不同领域的数据混合上进行训练,选择合适的混合比例可以显著提升下游性能。尽管关于数据混合和重新加权的文献众多,但现有工作大多是经验性的,尚不清楚辅助数据何时真正改进缩放律,而不仅仅是提供更多样本。为了深入理解这一问题,我们研究了一个高维混合数据回归模型,该模型具有共享的回归函数、异质协方差和噪声水平,以及可能以不同速率增长的数据集规模。我们在椭球参数约束下建立了通用协方差结构下的极小极大风险,并推导了交换协方差下岭回归测试误差的确定性等价形式。随后,我们专门研究了具有对齐幂律协方差谱的目标域和辅助域,此时理论给出了关于谱衰减、目标正则性和两个数据集相对增长的显式缩放律。这些缩放律识别了数据混合可证明地产生比单独使用任一数据集更快的缩放速率的机制。特别是,改进缩放律需要谱与域的相对样本量之间存在特定的相互作用。我们在语言模型上的数值实验展示了相同的定性现象:适当的数据混合比在任一域上单独训练能更快地降低目标域测试损失。

英文摘要

Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a high-dimensional mixed-data regression model with a shared regression function, heterogeneous covariances and noise levels, and dataset sizes that may grow at different rates. We establish the minimax risk under an ellipsoidal parameter constraint for the general covariance structure and derive deterministic equivalents for the test error of ridge regression under commutative covariances. We then specialize to a target domain and an auxiliary domain with aligned power-law covariance spectra, where the theory yields explicit scaling laws in terms of spectral decay, target regularity, and the relative growth of the two datasets. These laws identify regimes in which combining data mixtures provably yields a faster scaling rate than using either dataset alone. In particular, improving the scaling law requires a specific interplay between spectra and relative sample sizes of the domains. Our numerical experiments on language models exhibit the same qualitative phenomenon: appropriate data mixtures yield a faster decrease in target-domain test loss than training on either domain alone.

发表机构

  • Institute of Science and Technology Austria (ISTA)(奥地利科学技术研究所)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

↑