评估用于喷注标记的Transformer中的参数冗余
Assessing Parameter Redundancy in Transformers for Jet Tagging
AI总结:
本文针对Transformer喷注标记器参数冗余问题,通过引入沙漏结构和轻量级嵌入层构建变体,大幅减少参数并基本保持标记性能。
AI中文摘要:
基于Transformer的喷注标记器,如粒子Transformer(Particle Transformer,ParT)和多交互粒子Transformer(More-Interaction Particle Transformer,MIParT),通过利用喷注组分间的关联实现了出色的区分能力,但通常需要比早期深度学习标记器更多的可训练参数。本文研究能否用显著更少的参数实现相当的区分能力。我们引入一种沙漏结构,替换注意力模块中的前馈网络(FFNs),同时保留粒子交互注意力不变;还引入轻量级粒子嵌入层,替换原始的密集嵌入网络。将这两种修改应用于ParT和MIParT,分别得到沙漏(HG)变体ParT-HG和MIParT-HG。我们在顶夸克标记和夸克-胶子区分的基准数据集上评估这两种模型。两种变体均保持相当的标记性能,包括固定信号效率下的背景拒绝能力,且仅使用各自基线模型约48%和39.7%的参数。在更大的JetClass数据集上,准确率和AUC下降不足1%,且部分信号类别的背景拒绝能力略有下降。总体而言,我们的方法提供了减少Transformer喷注标记器参数数量的替代方案,同时在很大程度上保留其标记性能。
英文摘要:
Transformer-based jet taggers, such as the Particle Transformer (ParT) and the More-Interaction Particle Transformer (MIParT), achieve excellent discrimination by exploiting correlations among jet constituents, but often require more trainable parameters than earlier deep-learning taggers. In this paper, we investigate whether comparable discriminating power can be achieved with substantially fewer parameters. We introduce an hourglass structure that replaces the feed-forward networks (FFNs) in the attention blocks while leaving the particle-interaction attention unchanged. We also introduce a lightweight particle-embedding layer to replace the original dense embedding network. Applying both modifications to ParT and MIParT yields the hourglass (HG) variants ParT-HG and MIParT-HG, respectively. We evaluate both models on benchmark datasets for top tagging and quark-gluon discrimination. Both variants retain comparable tagging performance, including background rejection at fixed signal efficiencies, while using only approximately 48% and 39.7% of the parameters of their respective baselines. On the larger JetClass dataset, accuracy and AUC decrease by less than 1%, and background rejection also decreases for several signal classes. Overall, our approach provides an alternative way to reduce the parameter count of Transformer jet taggers while largely retaining their tagging performance.