发表机构
Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于 Python CCCL 的 Awkward Array CUDA 执行模型,消除自定义内核,支持操作融合与惰性优化,提升高能物理分析中 GPU 性能与可扩展性。
AI 中文摘要
Awkward Array 是高能物理(HEP)领域广泛使用的库,用于在 Python 中表示和操作嵌套的、可变长度的数据。之前的 CHEP 贡献已经探索了 Awkward Array 的 GPU 加速,证明了基于 CUDA 的后端的可行性和性能优势,同时也指出了与不规则数据访问、细粒度内核启动以及操作可组合性相关的局限性。在本贡献中,我们介绍了直接在先前工作基础上构建的最新进展,通过引入基于 Python CUDA 核心计算库(CCCL)的 Awkward Array CUDA 执行模型。使用 CCCL,我们消除了对自定义 CUDA 内核的需求,转而可以使用高级 Python 接口。基于 CCCL 的方法还支持将多个 Awkward 操作融合为数量减少的 CUDA 内核,解决了早期 GPU 实现中观察到的内核启动开销问题。惰性执行允许在生成内核之前构建和优化表达式图,从而提升涉及锯齿数组、组合操作和归约的分析工作流的性能。与早期方法相比,该设计还强调可扩展性,允许用户定义的 Python 代码以最少的样板代码并入 GPU 执行路径,且不破坏现有的分析语义。我们展示了性能研究,证明对于代表性的 HEP 分析模式,相比之前报告的即时 GPU 执行策略有所改进。这些进展将 Awkward Array 的 GPU 能力扩展至更可组合和可持续的后端,符合 HL-LHC 及以后基于 Python 分析的需求。
英文摘要
Awkward Array is a widely used library in high-energy physics (HEP) for representing and manipulating nested, variable-length data in Python. Previous CHEP contributions have explored GPU acceleration for Awkward Array, demonstrating the feasibility and performance benefits of CUDA-based backend while also identifying limitations related to irregular data access, fine-grained kernel launches, and composability of operations. In this contribution, we present recent developments that build directly on these earlier efforts by introducing a CUDA execution model for Awkward Array based on the Python CUDA Core Compute Libraries (CCCL). Using CCCL, we eliminate the need for custom CUDA kernels and can instead use a high-level Python interface. The CCCL-based approach also enables fusion of multiple Awkward operations into a reduced number of CUDA kernels, addressing kernel launch overhead observed in earlier GPU implementations. Lazy execution allows expression graphs to be constructed and optimized prior to kernel generation, improving performance for analysis workflows involving jagged arrays, combinatorial operations, and reductions. In contrast to earlier approaches, this design also emphasizes extensibility, allowing user-defined Python code to be incorporated into GPU execution paths with minimal boilerplate and without breaking existing analysis semantics. We present performance studies that demonstrate improvements over previously reported eager GPU execution strategies for representative HEP analysis patterns. These developments extend the GPU capabilities of Awkward Array toward a more composable and sustainable backend, aligned with the needs of Python-based analysis at the HL-LHC and beyond.
Comments8 pages, 8 figures, 28th CHEP (2026, Bangkok)