arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合专家粒子Transformer中的条件容量与路由

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

Kaushik Pendiyala, Haris Zia, Trevin Lee, Timothy Legge, Alejandro J. De Leon, Zihan Zhao, Aaron Wang, Abhijith Gandrakota, Jennifer Ngadiuba, Richard Cavanaugh, Javier Duarte

arXiv 2610.02701首次发表:更新:

发表机构

University of California San Diego; University of Illinois at Chicago; Fermi National Accelerator Laboratory; University of California Davis(加州大学圣迭戈分校; 伊利诺伊大学芝加哥分校; 费米国家加速器实验室; 加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在JetClass-II上对比密集与MoE粒子Transformer,发现top-1 MoE在计算量不变时提升精度,多专家激活增加计算成本但提升预测,并强调需区分存储容量、活跃计算、路由容量与组织。

AI 中文摘要

混合专家(MoE)模型可以在不按比例增加活跃计算量的情况下增加参数容量,但尚不清楚这种权衡在粒子物理Transformer中的表现。我们在包含188个类别的JetClass-II上研究了密集和MoE粒子Transformer,变化了专家数量、路由容量、top-K和辅助损失。我们发现,当避免token丢弃时,top-1 MoE模型在名义前向计算几乎不变的情况下优于密集基线,而进一步增加存储的专家数量仅带来很小的额外精度提升。每个token激活多个专家会在更高计算成本下带来额外的预测改进。路由分析显示,在某些配置中,专家分配与粒子身份和运动学特征的关联更强,但这种结构并不随分类性能单调增加。这些结果强调了在评估用于喷注分类的稀疏专家模型时,需要区分存储参数容量、活跃计算量、路由容量和路由组织。代码和实验配置可在该https URL获取。

英文摘要

Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly associated with particle identity and kinematics in some configurations, but this structure does not increase monotonically with classification performance. These results highlight the need to distinguish stored parameter capacity, active computation, routing capacity, and routing organization when evaluating sparse expert models for jet classification. Code and experiment configurations are available at https://github.com/kpendiyala/MPT.

Comments9 pages, 4 figures. Submitted to the ML4PS 2026 workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑