面向高多重度LHC过程的高效事例生成:基于归一化流的端到端GPU工作流
Efficient Event Generation for High-Multiplicity LHC Processes: An End-to-End GPU Workflow with Normalizing Flows
AI总结:
该研究提出一种集成归一化流与Pepper的端到端GPU事例生成工作流,针对LHC高多重度过程,实现十亿级未加权事例生成的端到端加速达两个数量级,缓解了蒙特卡洛统计瓶颈。
AI中文摘要:
为高多重度过程生成极大量未加权事例样本的工作受限于昂贵的矩阵元计算和较低的未加权效率。我们提出首个端到端GPU驻留的事例生成工作流,将归一化流(normalizing flows)提议与部分子水平事例生成器Pepper相集成。利用在线更新结合样本回放训练受螺旋度约束的耦合流,并将其部署在包含大量末态喷注的完整质子-质子碰撞过程的所有子过程中。该工作流中,基于Python的控制层与Pepper直接在设备内存中交换流生成的相空间点及对应的目标密度评估结果。控制层执行流采样、提议密度评估和未加权操作,而Pepper则评估定义目标密度的矩阵元、部分子分布函数(PDFs)和相空间因子,并以标准格式写入接受的事例。我们对比了特定子过程流(每个部分子子过程对应一个流)与分组条件流(在部分子内容相关的子过程间共享参数)的性能。该工作流针对$pp \ o e^+e^- + 4j$、$pp \ o e^+e^- + 5j$、$pp \ o t \ar t + 4j$、$pp \ o 4j$和$pp \ o 5j$产生过程进行基准测试。在四块H100 GPU上,我们为每个基准过程生成$10^9$个未加权事例。计入流训练成本后,该工作流相较于 standalone Pepper事例生成实现了最高达两个数量级的端到端加速,将数周的任务转变为亚日级计算。由此它使十亿级事例生成更具实用性,并为缓解高多重度对撞物理中蒙特卡洛统计瓶颈提供了途径。
英文摘要:
Producing very large unweighted event samples for high-multiplicity processes is limited by expensive matrix-element evaluations and low unweighting efficiencies. We present the first end-to-end GPU-resident event-generation workflow that integrates normalizing-flow proposals with the parton-level event generator Pepper. Helicity-conditioned coupling flows are trained using online updates supplemented by sample replay and deployed across all subprocesses of complete proton--proton collision processes with many final-state jets. In this workflow, a Python-based control layer and Pepper exchange flow-generated phase-space points and the corresponding target-density evaluations directly in device memory. The control layer performs flow sampling, proposal-density evaluation, and unweighting, while Pepper evaluates the matrix elements, PDFs, and phase-space factors defining the target density and writes the accepted events in standard formats. We compare subprocess-specific flows, with one flow per partonic subprocess, to grouped conditional flows that share parameters among subprocesses with related parton content. The workflow is benchmarked for $pp \to e^+e^- + 4j$, $pp \to e^+e^- + 5j$, $pp \to t \bar t + 4j$, $pp \to 4j$, and $pp \to 5j$ production. On four H100 GPUs, we generate $10^9$ unweighted events for each benchmark process. Including the cost of flow training, the workflow achieves end-to-end speedups of up to two orders of magnitude over standalone Pepper event generation and turns a multi-week task into a sub-day computation. It thereby makes billion-event production more practical and offers a pathway to alleviating the Monte Carlo statistics bottleneck in high-multiplicity collider physics.