缓解GPU光学光子蒙特卡洛中的warp分歧:通过着色器执行重排序实现一个数量级的加速
Mitigating warp divergence in GPU optical photon Monte Carlo: an order of magnitude speedup with Shader Execution Reordering
浏览论文内容
中文总结 AI 辅助
该研究针对GPU光学光子蒙特卡洛的warp分歧问题,采用NVIDIA SER技术重排光子分组,在液态氩时间投影室基准测试中实现约15倍端到端加速,同时降低GPU能耗约90%。
中文摘要 AI 辅助
在高能物理、核物理和医学物理领域,光学光子输运蒙特卡洛方法是探测器模拟中常见的性能瓶颈。GPU光线追踪可加速光子传播,但每个光子对应一个线程的兆核函数会因光子寿命差异大而出现SIMT执行分歧:短寿命光子会使部分线程束(warp)的通道处于非活跃状态,而少数长寿命光子则会延迟整个warp的完成。我们在基于OptiX的光学光子输运核中评估了NVIDIA Shader Execution Reordering(SER,着色器执行重排序),并利用它在传播过程中对存活光子进行重新分组。在一个包含约6100万光子、来自2.5 GeV电磁簇射的14.7千吨液态氩时间投影室基准测试中,启用SER的执行使传播核的周期减少了16.55(23)倍,端到端光学模拟时间(含数据传输和初始化)减少了14.72(30)倍,且输出击中结果完全一致。我们将加速比拆解为:启用SER时引入的启动配置变化贡献了3.2倍,执行重排序本身贡献了额外的5.1倍。性能分析显示活跃通道恢复是主要加速机制,而分支效率和内存带宽变化很小。体积相干SER提示可在 bulk氩中启用分析型导航快速路径,额外带来1.13倍的墙钟时间加速。利用SER使每个模拟事件的GPU能耗降低了约90%。
英文摘要
In high energy, nuclear, and medical physics, optical photon transport Monte Carlo is a frequent bottleneck in detector simulation. GPU ray tracing accelerates photon propagation, but one thread per photon megakernels suffer from SIMT execution divergence when photon lifetimes vary widely. Short-lived photons leave inactive lanes while a few long-lived photons delay warp completion. We evaluate NVIDIA Shader Execution Reordering (SER) in an OptiX based optical photon transport kernel and use it to regroup surviving photons during propagation. In a 14.7 kton liquid argon time projection chamber benchmark with ~61 million photons from a 2.5 GeV electromagnetic shower, SER capable execution reduces propagation kernel cycles by 16.55(23)x and end-to-end optical simulation time, including transfers and initialization, by 14.72(30)x with bit-identical hit output. We decouple the speed-up into a 3.2x contribution from the launch configuration changes introduced when SER is enabled and a further 5.1x from executing the reorder. Profiling identifies active lane recovery as the dominant mechanism, while branch efficiency and memory bandwidth change little. A volume coherent SER hint enables an analytic navigation fast path in bulk argon, adding a further 1.13x wall-time speedup. Utilizing SER resulted in a ~90% reduction in GPU energy per simulated event.