arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPUPHOT:一个用于高性能GPU加速测光与分布式天文数据归算的Python框架

GPUPHOT: A Python Framework for High-Performance GPU-Accelerated Photometry and Distributed Astronomical Data Reduction

Samuel Lemes-Perera, Miguel R. Alarcon, Miquel Serra-Ricart, Pino Caballero-Gil

arXiv 2609.32375首次发表:更新:

发表机构

Light Bridges S.L.; Universidad de La Laguna (ULL); Instituto Tecnológico y de Energías Renovables (ITER); Instituto de Astrofísica de Canarias (IAC)(Light Bridges S.L.; 拉帕尔马大学; 技术与可再生能源研究所; 加那利天体物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GPUPHOT是一个开源Python框架,利用GPU加速实现实时测光和天体测量,在多种GPU平台上高效处理大视场图像,显著提升源检测速度,并确保结果可复现,适用于自主巡天设施。

AI 中文摘要

我们介绍GPUPHOT,一个开源的Python框架,用于天文CCD和科学CMOS图像的GPU加速实时测光和天体测量。其七个阶段中有四个完全在GPU上通过CuPy运行(背景估计、源检测、PSF建模和孔径测光);星表交叉匹配根据问题规模选择CPU或GPU后端,而零点定标和天体测量定标在CPU上运行,位于容器化的分布式系统内。基于检查点的GPU内存监控和可选的空间分箱回退机制,使单个容器镜像能在从4 GB笔记本电脑显卡到80 GB数据中心加速器的GPU上运行。我们在来自两个机器人设施(两米双筒望远镜和瞬变巡天望远镜)的19次观测上对框架进行基准测试,图像大小从4.2到151.2百万像素,包含228到134,206个已编目源,在八个NVIDIA GPU平台上以每图像一个进程的方式运行。H100在31.9秒内端到端完成最密集的151.2百万像素视场,而A100上的GPU源检测比同一帧上的CPU库sep快6-15倍。一旦确定了三个非确定性来源,星表和零点在不同GPU型号和主机CPU上的一致性达到小数点后第四位。在这种每图像一个进程的模式下,CuPy 14 / cuML堆栈相比CuPy 12显示出7%更低的峰值内存和38%更高的每图像延迟中位数,两者均归因于RAPIDS分配器在导入时取代了CuPy内存池;恢复内存池消除了内存差异,并将数据中心GPU上的中位数开销降至11%,输出结果相同。长期运行的worker(生产模式)使两个堆栈的性能接近持平。我们推荐CuPy 14堆栈,因其可复现性和维护良好的库,而非内存或速度优势。GPUPHOT及其基准测试脚本已发布,供需要以巡天节奏进行端到端测光归算的自主设施使用。

英文摘要

We present GPUPHOT, an open-source Python framework for GPU-accelerated real-time photometry and astrometry of astronomical CCD and scientific CMOS images. Four of its seven stages run entirely on the GPU through CuPy (background estimation, source detection, PSF modeling and aperture photometry); the catalog crossmatch selects a CPU or GPU backend by problem size, and the zero-point and astrometric calibrations run on the CPU, within a containerized distributed system. Checkpoint-based GPU memory monitoring and an opt-in spatial-binning fallback let a single container image run on GPUs from a 4 GB laptop card to 80 GB datacenter accelerators. We benchmark the framework on nineteen observations from two robotic facilities, the Two-meter Twin Telescope and the Transient Survey Telescope, spanning 4.2 to 151.2 megapixels and 228 to 134,206 cataloged sources, on eight NVIDIA GPU platforms with one process per image. The H100 completes the densest 151.2-megapixel field end-to-end in 31.9 s, and GPU source detection on the A100 is 6-15x faster than the CPU library sep on the same frames. Once three sources of non-determinism are pinned, catalogs and zero points agree to the fourth decimal across GPU models and host CPUs. In this one-process-per-image regime the CuPy 14 / cuML stack showed a 7% lower peak memory and a median 38% higher per-image latency than CuPy 12, both due to the RAPIDS allocator displacing the CuPy memory pool at import; restoring the pool removes the memory difference and cuts the median overhead on datacenter GPUs to 11%, with identical output. Long-lived workers, the production mode, run the two stacks near parity. We recommend the CuPy 14 stack for its reproducibility and maintained libraries, not for memory or speed. GPUPHOT and its benchmark scripts are released for autonomous facilities that require end-to-end photometric reduction at survey cadence.

Comments39 pages, 8 figures. Submitted to Astronomy and Computing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑