AI 中文总结
本文针对高能物理领域的不规则工作负载加速难题,推进Awkward Array GPU后端开发,通过CUDA实现等优化提升了GPU处理性能,并完成了多实现的基准测试对比。
AI 中文摘要
Awkward Array是一款用于表示和处理嵌套变长数据的Python库,在高能物理领域应用广泛。随着HL-LHC分析愈发依赖加速器硬件,不规则工作负载的高效执行已成为关键需求。密集数值数组可自然适配GPU,但嵌套变长数据结构因需间接索引、分段操作及不规则内存访问模式,加速难度显著更高。本文介绍Awkward Array GPU后端的最新进展,包括基于NVIDIA CUDA核心计算库(CCCL)构建的CUDA实现、优化的内存管理,以及针对不规则数组的分段归约算法。这些进展在保留现有Python编程模型的同时,大幅提升了GPU对不规则工作负载的吞吐量。文中还描述了后端架构、自动验证框架,以及对比CPU、CuPy和CUDA实现的基准测试结果。
英文摘要
Awkward Array is a Python library for representing and processing nested, variable-length data that is widely used in high-energy physics. As HL-LHC analyses increasingly rely on accelerator hardware, efficient execution of irregular workloads has become essential. While dense numerical arrays map naturally to GPUs, nested and variable-length data structures remain significantly more difficult to accelerate because they require indirect indexing, segmented operations, and irregular memory access patterns. We present recent developments in the Awkward Array GPU backend, including CUDA implementations built on NVIDIA CUDA Core Compute Libraries (CCCL), optimized memory management, and segmented reduction algorithms for ragged arrays. These developments preserve the existing Python programming model while substantially improving GPU throughput on irregular workloads. We describe the backend architecture, automated validation framework, and benchmark results comparing CPU, CuPy, and CUDA implementations.