arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优化从红移概率分布函数中提取信息的方法

Optimizing the extraction of information from redshift probability distribution functions

Rodrigo Duarte, Valerio Marra

arXiv 2607.26822首次发表:更新:

AI 中文总结

该研究提出turboPDZ机器学习框架,从HSC-SSP PDR3的PDZ中优化测光红移的点估计与可靠性度量,性能优于星表指标,可高效识别不可靠天体。

AI 中文摘要

测光红移是大尺度结构分析的核心要素,但从测光红移(photo-$z$)管道输出的红移概率分布函数(PDZ)中提取最优点估计值与可靠性度量仍是未解决的挑战。我们提出turboPDZ,一种直接从PDZ优化上述两个量的机器学习框架。将该框架应用于广域层(Wide)和深场层(DUD)中三个独立的HSC-SSP PDR3管道(DEmP、DNNz、Mizuki)输出的PDZ:每个PDZ通过主成分分析(PCA)压缩后与汇总描述符结合;用Optuna在复合目标下优化的多层感知器生成优化点估计值$z_{\rm ml}$;第二个在对数空间训练并校准的网络输出不确定性$\tau_{\rm ml}$,据此通过百分位排名得到可靠性评分$r_{\rm ml}$。$z_{\rm ml}$在全部6种管道-层组合的归一化绝对偏差中值($\tau_{\rm NMAD}$)和$\tau_{0.15}$指标上均优于星表的最优红移$z_{\rm best}$。以$\tau_{\rm NMAD}$和$\tau_{0.15}$随保留天体占比的曲线下面积衡量,$r_{\rm ml}$筛选星系的效率高于星表的风险和置信度指标。对于模板拟合管道Mizuki,星表指标失效严重,其曲线下面积(AUC)可达$r_{\rm ml}$的10倍,而$r_{\rm ml}$能在所有红移区间正确识别不可靠天体。特征重要性分析显示互补模式:点估计由PCA分量和位置统计主导,可靠性估计则依赖PCA分量和峰值统计。该管道与巡天无关,公开可用,训练模型及优化后量作为增值星表发布。

英文摘要

Photometric redshifts are essential for large-scale structure analyses, yet extracting optimal point estimates and reliability measures from the probability distribution functions (PDZs) delivered by photo-$z$ pipelines remains an open challenge. We introduce turboPDZ, a machine-learning framework that optimizes both quantities directly from the PDZ. We apply the framework to PDZs from the three independent HSC-SSP PDR3 pipelines (DEmP, DNNz, Mizuki) across Wide and DUD layers. Each PDZ is compressed via PCA and combined with summary descriptors; a multilayer perceptron, optimized with Optuna under a composite objective, produces the optimized point estimate $z_{\rm ml}$. A second network, trained in log-space and calibrated, yields the uncertainty $σ_{\rm ml}$, from which the reliability score $r_{\rm ml}$ is derived via percentile ranking. $z_{\rm ml}$ outperforms the catalog $z_{\rm best}$ in $σ_{\rm NMAD}$ and $η_{0.15}$ across all six pipeline-layer combinations. $r_{\rm ml}$ filters galaxies more efficiently than the catalog risk and confidence indicators, as measured by the area under the $σ_{\rm NMAD}$ and $η_{0.15}$ versus retained-fraction curves. For Mizuki, the template-fitting pipeline, the catalog indicators fail dramatically, with AUC values up to ten times larger than those of $r_{\rm ml}$, whereas $r_{\rm ml}$ correctly identifies unreliable objects across all redshift regimes. Feature-importance analysis reveals complementary patterns: point estimation is dominated by PCA components and location statistics, while reliability estimation depends on PCA components and peak statistics. The pipeline is survey-independent, publicly available at https://github.com/valerio-marra/turboPDZ, and trained models plus optimized quantities are released as a value-added catalog.

Comments14 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑