arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23377hep-excs.AI

训练前预测:粒子物理基础模型的缩放定律

Predict before you train: Scaling Laws for particle physics foundation models

Jan-Lucas Uslu, Benjamin Nachman, Christopher Re

首次发表
浏览论文内容

中文总结 AI 辅助

研究粒子物理基础模型训练成本高及缩放回报难测问题,通过在小模型上拟合联合缩放定律预测大模型损失,还将其与下游物理性能关联,发布相关预训练模型、训练方法及代码。

中文摘要 AI 辅助

粒子物理学中最大的机器学习模型训练成本也最高,且在计算投入前无法估计给定架构的缩放回报。虽已为喷注拟合了缩放定律,但尚未证明其能预测未在其上拟合的模型性能。我们表明,对于在对撞机喷注上预训练的通用变压器,可以进行预测。仅在小模型上拟合联合模型和数据缩放定律,跨越三个数量级的训练计算量,我们能将之后用超过百倍计算量训练的模型损失预测在百分之一以内。然后将预测与下游物理性能联系起来:在两个标准标记基准上,预训练损失越低,微调损失越低,微调后背景拒绝率越高。在此模型家族和这些任务中,在训练任何大型模型之前,计算预算可转化为预期物理性能。最终前沿模型在准确性、AUC和夸克/胶子拒绝率方面与在同一语料库上训练的当前最先进物理感知基础模型的已发表数据一致,仅在顶级标记的高纯度尾部,物理感知模型有残余优势。我们发布了五个跨多种大小的预训练模型,以及完整的训练方法和代码。

英文摘要

The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on. We show that, for a generic transformer pretrained on collider jets, it can be forecast. Fitting a joint model-and-data scaling law on small models alone, spanning three orders of magnitude of training compute, we predict the loss of models trained afterward with more than one hundred times more compute to within one percent. We then connect the forecast to downstream physics performance: across two standard tagging benchmarks, lower pretraining loss yields systematically lower fine-tuning loss and higher background rejection after fine-tuning. Within this model family and these tasks, a compute budget can therefore be translated into expected physics performance before any large model is trained. The final frontier model is consistent with the published numbers for current state-of-the-art physics-aware foundation models trained on the same corpus, on accuracy, AUC, and quark/gluon rejection, with a residual edge for the physics-aware model only in the high-purity tail of top tagging. We release five pretrained models spanning multiple sizes, together with the complete training recipe and code.

发表机构

  • Stanford University(斯坦福大学)
  • SLAC National Accelerator Laboratory(SLAC国家加速器实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑