发表机构
Technical University of Munich; SLAC National Accelerator Laboratory; University College London; University of Geneva; Humboldt-Universität zu Berlin; NSF AI Institute for Artificial Intelligence and Fundamental Interactions; Munich Center for Machine Learning(慕尼黑工业大学; SLAC国家加速器实验室; 伦敦大学学院; 日内瓦大学; 柏林洪堡大学; NSF人工智能与基础相互作用人工智能研究所; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出系统化流程以推导稳健缩放定律并比较HEP任务设计选择,在110亿喷注数据集上验证,首次预测联合最优超参数,并发现辅助目标与低层级输入可提升性能。
AI 中文摘要
语言模型等机器学习领域近期的大部分进展,来自于预测性能随训练投入变化的缩放定律。在高能物理(HEP)中,现已观察到类似行为。为促进进一步研究,我们提出了一套系统化流程,用于推导稳健的缩放定律,并在HEP任务的相关预算轴上比较设计选择。我们首先在玩具问题上验证完整的缩放轨迹,然后将该流程应用于ATLAS JetSet2数据集(约110亿喷注)上的多任务Transformer,涵盖计算受限和数据受限两种情形。对于后者,我们首次(据我们所知)预测了在早停条件下的联合最优模型规模、训练时长、学习率和批大小。在计算最优缩放下,我们恢复了模型与数据集规模近相等的$\sqrt{C}$依赖关系,并发现辅助目标在相同计算预算下降低了主喷注分类损失。将输入扩展到更低层级的数据,系统性地降低了损失,同时缩放指数几乎不变。幂律区域的起始本身由规模决定:当数据集规模低于阈值时,损失对高计算量缩放的信息量很小,这凸显了大规模、高质量全模拟数据集作为缩放研究及HEP基础模型开发基础的价值。
英文摘要
Much of the recent progress in machine learning domains such as language models has come from scaling laws that predict performance as a function of training effort. In high-energy physics (HEP) similar behavior has now been observed. To aid further study, we present a systematic procedure to derive robust scaling laws and compare design choices on the relevant budget axes for HEP tasks. We first validate the full scaling trajectory on toy problems and then apply the procedure to multi-task transformers on the ~11 billion-jet ATLAS JetSet2 dataset, in both the compute- and data-constrained regimes. For the latter, we predict, to the best of our knowledge for the first time, the jointly optimal model size, training horizon, learning rate and batch size under early stopping. At compute-optimal scaling, we recover a near-equal $\sqrt{C}$ dependence of model and dataset size, and find that auxiliary objectives lower the primary jet-classification loss at equal compute budget. Expanding the inputs toward lower-level data systematically lowers the loss while leaving the scaling exponent nearly unchanged. The onset of the power-law regime is itself set by scale: below a threshold in dataset size the loss carries little information about high-compute scaling, underscoring the value of large, high-quality full-simulation datasets as a foundation for scaling studies and the development of foundation models in HEP.
Comments40 pages, 60 figures, 7 tables