AI 中文总结
本研究针对JUNO的OMILREC重建算法,通过分阶段等效保持优化实现了最高8.6倍的单线程加速,保持物理输出一致,为大型探测器似然重建加速提供了可转移模板。
AI 中文摘要
江门地下中微子天文台(JUNO)采用OMILREC进行每个事例的顶点与能量重建,OMILREC是一种最大似然拟合方法,在每次Minuit函数评估中扫描全部17612个大型光电倍增管(LPMT),每个事例约进行470次。该内层循环是重建CPU成本的主要来源。性能分析显示,生产算法受延迟限制,由于虚函数调度、ROOT直方图指针追逐和重复计算,仅达到标量浮点峰值的9.9%。我们应用分阶段的“等效保持”优化:扁平化数据布局、可向量化几何、Minuit不变工作的提升、每个事例的预计算、拟合阶段的循环拆分与索引,以及低精度快速路径。每个阶段均与未修改代码的冻结参考进行核对。优化后的实现,在英特尔至强铂金8358P上单线程加速8.06倍(从1524.8毫秒/事例降至189.2毫秒/事例),在AMD EPYC 9654上单线程加速5.22倍(从705.1毫秒/事例降至134.9毫秒/事例),进一步优化后加速提升至8.6倍(177.7毫秒/事例)。前七个版本的似然保持位级一致,后续相对漂移在1.3×10⁻¹⁴内,低于10⁻¹³的约定。对于典型事例,重建顶点与能量与基线的偏差在4毫米和7千电子伏以内;少数边界情况因优化的最小化器种子达到不同的有效最小值。约861000个⁶⁸Ge校准事例的八项物理接受度测试也通过。该工作流程在遵守验证门的AI编码代理协助下开发,为大型中微子和对撞机探测器中基于似然的重建加速提供了可转移模板,且不改变物理输出。
英文摘要
The Jiangmen Underground Neutrino Observatory (JUNO) reconstructs each event's vertex and energy with OMILREC, a maximum-likelihood fit that scans all $17{,}612$ large photomultiplier tubes (LPMTs) in every Minuit function evaluation, about $470$ times per event. This inner loop dominates the reconstruction CPU cost. Profiling shows that the production algorithm is latency-bound, sustaining only $9.9%$ of scalar floating-point peak because of virtual-function dispatch, ROOT-histogram pointer chasing, and repeated computation. We apply staged \emph{equivalence-preserving} optimizations: flattened data layouts, vectorizable geometry, hoisting of Minuit-invariant work, per-event precomputation, fit-phase loop splitting and indexing, and reduced-precision fast paths. Each stage is checked against a frozen reference from the unmodified code. The optimized implementation achieves single-thread speedups of $8.06\times$ ($1524.8 \rightarrow 189.2$~ms/event) on an Intel Xeon Platinum~8358P and $5.22\times$ ($705.1 \rightarrow 134.9$~ms/event) on an AMD~EPYC~9654, increasing to $8.6\times$ ($177.7$~ms/event) after further optimization. The likelihood remains bit-identical through the first seven releases and later agrees within a relative drift of $1.3\times10^{-14}$, below the $10^{-13}$ contract. For typical events, reconstructed vertex and energy agree with the baseline within $4$~mm and $7$~keV; a few boundary cases reach different valid minima owing to an improved minimizer seed. An eight-metric physics-acceptance test also passes on about $861{,}000$ $^{68}$Ge calibration events. Developed with assistance from an AI coding agent operating under these verification gates, this workflow offers a transferable template for accelerating likelihood-based reconstruction in large neutrino and collider detectors without changing physics output.
Comments7 pages, 1 figure