基于归一化流与并行化的脉冲星计时计算加速研究
On the Acceleration of Pulsar Timing computations using Normalising Flows and Parallelisation
- Physical Research Laboratory(物理研究实验室)
- Ahmedabad University(艾哈迈达巴德大学)
- IISER Bhopal(印度科学教育研究所博帕尔分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对脉冲星计时计算的高维度等瓶颈,采用POCOMC采样技术与并行化方案,对比了不同软件包的加速效果,为PTA分析提供了高效计算途径。
AI中文摘要:
单脉冲星噪声分析(SPNA)和脉冲星计时阵列(PTA)数据集上的引力波(GW)搜索长期面临计算瓶颈,该瓶颈源于PTA似然景观的高维度、多模态,以及各类单脉冲星噪声与集合级公共噪声过程间的强相关性。我们通过采用在POCOMC软件包中实现的基于归一化流的预处理蒙特卡洛采样技术,首次将其应用于PTA特定计算,并将实现的加速效果与广泛使用的PTMCMCSAMPLER和DYNESTY软件包进行比较。我们进一步通过在高性能计算(HPC)资源上利用PARALLEL_BILBY架构搭配DYNESTY,研究了在不断增加的通信节点阵列上通过并行化实现的加速效果。我们在包含SPNA和公共红噪声(CRN)分析的真实长基线模拟数据集上测试了加速效果。我们发现PARALLEL_BILBY在并行化方面效率最高,使用16个节点进行空间不相关和Hellings and Downs相关CRN搜索时,运行时间分别约为10分钟和100分钟;POCOMC在单节点性能上表现更优,相关搜索仅需约10小时;PTMCMCSAMPLER效率最低。我们预计POCOMC对PTA分析具有重要意义,无需任何GPU或HPC支持,即可在可管理的时间跨度内执行集合级引力波搜索。这些结果对数据量不断增长且需纳入更复杂模型的情况具有长远影响,否则这些模型会因相关计算成本过高而无法实现。
英文摘要:
Single-Pulsar Noise Analysis (SPNA) and Gravitational Wave (GW) searches done on Pulsar Timing Array (PTA) datasets have everlastingly suffered from the computational bottleneck arising due to high dimensionality and multi-modality of the PTA likelihood landscape, along with strong correlations amongst various single-pulsar noises and ensemble-level common noise processes. We addressed this outstanding issue by employing a Normalising Flow-based Preconditioned Monte-Carlo sampling technique implemented in the POCOMC package, for the first time on PTA-specific computations, and comparing the achieved acceleration with the widely used PTMCMCSAMPLER and DYNESTY packages. We further investigated the acceleration achieved via parallelisation over an increasing array of communicating nodes on a high-performance computing (HPC) resource, by employing the PARALLEL_BILBY architecture with DYNESTY. We tested the acceleration on realistic long baseline simulated datasets with SPNA and Common Red Noise (CRN) analysis. We found PARALLEL_BILBY to be the most efficient in parallelisation, achieving a runtime of ~10min and ~100min with 16 nodes for spatially uncorrelated and Hellings and Downs-correlated CRN searches, respectively. POCOMC outperforms in single node performance requiring only ~10h for correlated search. PTMCMCSAMPLER was found to be the least efficient. We envisage POCOMC to be of great importance for PTA analyses, without requiring any GPU or HPC support, while also performing ensemble-level GW searches within a manageable time span. These results have everlasting implications with increasing data volumes and need to incorporate more complicated models, which were otherwise beyond reach due to the associated computational costs.