面向粒子物理硬件加速器的基础模型
Towards Foundation Models on Hardware Accelerators for Particle Physics
- Stanford University(斯坦福大学)
- SLAC National Accelerator Laboratory(SLAC国家加速器实验室)
- University of California, Irvine(加州大学尔湾分校)
- Nagoya University(名古屋大学)
- Kobayashi-Maskawa Institute(小林诚益川敏英研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对粒子物理硬件加速器部署,通过知识蒸馏将OmniLearned基础模型迁移至无注意力Deep Sets网络,三种改进策略在低信号效率背景抑制中增益最大。
AI中文摘要:
带宽限制要求许多粒子物理实验在定制硬件或固件上运行简化算法以做出实时决策,这与离线分析不同,在离线分析中延迟通常不是限制因素。例如,在对撞机的粒子喷注标记任务中,最先进的性能由具有数亿参数、在数十亿喷注上预训练的基础模型实现。我们使用知识蒸馏将此类模型所学内容迁移到高效网络中,以部署到硬件加速器。教师模型是在顶夸克喷注标记上微调的OmniLearned基础模型;学生模型是无注意力机制的Deep Sets网络。我们展示了学生模型性能提升的三种方式:在Deep Sets架构中添加消息传递层、在教师模型的软标签而非仅真实标签上训练、以及从预训练教师模型而非从头训练的同架构模型中蒸馏。在每种情况下,增益在低信号效率下的背景抑制中最大,而这正是对触发系统最相关的场景。
英文摘要:
Bandwidth constraints require many particle physics experiments to make real-time decisions on custom hardware or firmware running simplified algorithms, unlike offline analysis, where latency is usually not a limiting factor. For example, for particle jet tagging at colliders, state-of-the-art performance is achieved by foundation models with hundreds of millions of parameters pre-trained with billions of jets. We use knowledge distillation to transfer what such models have learned into efficient networks towards deployment in hardware accelerators. The teacher is the OmniLearned foundation model fine-tuned on top quark jet tagging; the student is an attention-free Deep Sets network. We demonstrate three ways the student's performance improves: adding a message-passing layer to the Deep Sets architecture, training on the teacher's soft labels rather than on ground-truth labels alone, and distilling from a pretrained teacher rather than from the same architecture trained from scratch. In each case the gain is largest in the background rejection at low signal efficiency, the regime that is most relevant for a trigger.