CatBoost与光谱能量分布拟合:在受控测光不完备性下估计星系属性
CatBoost versus Spectral Energy Distribution-Fitting: Estimating Galaxy Properties under Controlled Photometric Incompleteness
浏览论文内容
中文总结 AI 辅助
本文对比CatBoost与理想SED拟合在受控测光不完备下估计星系属性的性能,发现CatBoost在高缺失率下仍表现良好,且在质量、SFR估计上优于SED拟合,可用于处理不完美的天文测光数据。
中文摘要 AI 辅助
从测光数据估计星系物理参数时,天文巡天普遍存在的缺失测量值构成了根本性挑战。本文利用Horizon-AGN流体动力学模拟生成的 mock 星表(提供真实物理参数),对原生处理缺失数据的梯度提升算法CatBoost进行评估,评估场景刻意具有对抗性:我们在12个测光波段上训练CatBoost,每个波段注入的缺失程度逐步提升(每波段缺失率为10%、20%、30%),并将其性能与理想化的参数化光谱能量分布(SED)拟合基准进行对比,该基准使用完整的26波段测光数据(无缺失,所有波段可用)。这种非对称设计代表了传统方法的上限、最佳情况基线。尽管存在这种刻意的劣势,CatBoost的性能下降仍有限:从完整数据过渡到30%缺失时,质量的均方根误差(RMSE)从0.08 dex升至0.18 dex,恒星形成率(SFR)的RMSE从0.41 dex升至0.53 dex,红移的RMSE从0.20升至0.28,偏差仍接近零。与理想SED拟合基准相比,在仅12个波段且缺失率30%的情况下训练的CatBoost,在质量上实现了更低的误差(0.18 vs. 0.28 dex),在SFR上也实现了更低的误差(0.53 vs. 0.57 dex),并消除了SED拟合结果中存在的系统偏差;对于红移,SED拟合基准的归一化绝对偏差(NMAD)更低(极端缺失时为0.030 vs. 0.093),而CatBoost保持了更小的偏差。这些结果表明,至少在本文探索的条件下,CatBoost对缺失值的原生处理可为从不完美的测光巡天中提取星系属性提供实用优势。
英文摘要
Estimating galaxy physical parameters from photometric data is fundamentally challenged by missing measurements that are endemic to astronomical surveys. Using a mock catalog from the Horizon-AGN hydrodynamical simulation that provides the true physical parameters, we evaluate CatBoost, a gradient-boosting algorithm that natively handles missing data, under a deliberately adversarial scenario: we train it on 12 photometric bands with increasing levels of injected missingness (10%, 20%, 30% missing per band) and compare its performance against an idealised parametric spectral energy distribution (SED)-fitting reference that uses complete 26-band photometry (no missing data, all bands available). This asymmetric design represents an upper-bound, best-case baseline for traditional methods. Despite this intentional disadvantage, CatBoost's performance degradation is limited: moving from complete data to 30% missing, mass RMSE increases from 0.08 to 0.18 dex, SFR RMSE from 0.41 to 0.53 dex, and redshift RMSE from 0.20 to 0.28, while bias remains near zero. Against the ideal SED-fitting reference, CatBoost trained on only 12 bands with 30% missing values achieves lower errors for mass (0.18 vs. 0.28 dex) and SFR (0.53 vs. 0.57 dex) and removes the systematic biases present in the SED-fitting results. For redshift, the SED-fitting reference has lower NMAD (0.030 vs. 0.093 at extreme missingness) while CatBoost maintains smaller bias. These results suggest that CatBoost's native handling of missing values can offer practical advantages for extracting galaxy properties from imperfect photometric surveys, at least under the conditions explored here.