发表机构
Universitat de València; Università di Bologna; Istituto Nazionale di Astrofisica(瓦伦西亚大学; 博洛尼亚大学; 意大利国家天体物理研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 cosmokdtree,一个基于 Fortran 和 OpenMP 的灵活 k-d 树库,支持任意维度,性能优于 CPU 替代方案,并应用于天体物理中的聚类与插值问题。
AI 中文摘要
现代数值天体物理学应用普遍需要高效的空间查询技术,例如对数十亿个分辨率元素进行邻居搜索或密度估计。为满足对这些任务量身定制的快速且多功能工具的日益增长的需求,我们提出了 cosmokdtree,这是一个用 Fortran 编写的快速、灵活的多用途 k-d 树实现,并通过 OpenMP 指令实现并行化。我们的库支持任意维度和空间分布,专为中等规模应用(性能测试高达 10^9 个点)的高效树构建和快速查询性能而设计。还提供了与该模块耦合的 Python 绑定。所有这些要素共同构成了一个适用于分析目的的 k-d 树软件包。我们在多种场景下对 cosmokdtree 进行了基准测试,包括不同的点分布、维度和并行扩展。我们还将它的性能与广泛使用的 scipy 实现以及其他高效替代方案(如 coretran 和基于 GPU 的版本)进行了比较,结果表明 cosmokdtree 在保持合理内存使用的同时,始终比 CPU 替代方案实现更低的构建和查询时间。我们的树构建阶段实现,在中端工作站级 CPU 上执行,其性能接近在高端消费级显卡上运行的基于 GPU 的实现,尽管数据中心 GPU(不在我们的比较范围内)仍可能提供显著更高的性能,因此无法得出更广泛的结论。我们进一步展示了一些解决天体物理学中常见问题的应用,即 friends-of-friends 聚类算法和粒子到网格赋值过程。该代码已公开发布,旨在作为计算应用(尤其是在天体物理学中)的灵活多用途工具。
英文摘要
Modern numerical astrophysics applications present a common demand for efficient spatial querying techniques, such as neighbor searches or density estimations over billions of resolution elements. To address the growing need for fast and versatile tools tailored to these tasks, we present cosmokdtree, a fast and flexible multi-purpose k-d tree implementation in Fortran, parallelized with OpenMP directives. Our library supports arbitrary dimensionality and spatial distributions, and is designed for efficient tree construction and fast query performance for moderate-scale applications (performance tested up to 10^9 points). Python bindings coupled to the module are also provided. All these ingredients yield a k-d tree package suitable for analysis purposes. We benchmark cosmokdtree across a variety of scenarios, including different point distributions, dimensionality, and parallel scaling. We also compare its performance against the widely used scipy implementation and other efficient alternatives such as coretran's and a GPU-based version, showing that cosmokdtree consistently achieves lower construction and query times than CPU alternatives while keeping a reasonable memory usage. Our tree-building phase implementation, executed on a mid-range workstation-class CPU, approaches the performance of GPU-based implementations when run on high-end consumer graphic cards, although data-center GPUs, out of the scope for our comparison, could still deliver substantially higher performance and a broader conclusion cannot be extracted. We further present some applications tackling common problems in astrophysics, namely, the friends-of-friends clustering algorithm and the particle-to-mesh assignment process. The code is publicly released and intended to serve as a flexible multi-purpose tool for computational applications in a wide range of scenarios, particularly in astrophysics.
Comments29 pages, 15 figures; published in Astronomy & Computing
Journal refA&C, 57, 101175 (2026)
DOI:10.1016/j.ascom.2026.101175