用于巡天规模活动星系核候选体优先级排序的六类BPT星系分类:深度表格模型与信息性缺失信号
Six-Class BPT Galaxy Classification for Survey-Scale AGN Candidate Prioritization: Deep Tabular Model and Informative Missingness Signals
浏览论文内容
中文总结 AI 辅助
研究利用机器学习模型,以测量量、导出线比值和缺失数据模式为输入,分析147万个星系,测试能否重现六类BPT标签。通过多种架构对比及评估,发现卷积神经网络-Transformer性能最强,缺失指标有预测信息,模型可作AGN候选体排名工具并补充传统诊断。
中文摘要 AI 辅助
鲍德温-菲利普斯-特勒维奇(BPT)图被广泛用于将星系分类为恒星形成系统、复合星系和活动星系核(AGN),但其巡天规模应用受限于对高信噪比发射线测量的要求。我们测试机器学习模型能否在使用测量量、导出的线比值和潜在信息丰富的缺失数据模式作为输入时重现六类BPT标签。我们用一个27维特征集分析了147万个星系,该特征集结合了原始巡天测量、导出量和缺失指标。将五种深度表格架构与梯度提升树和经典机器学习基线进行基准测试,并通过硬分类指标、精确召回曲线、top-k检索、消融测试和特征解释诊断来评估结果概率。卷积神经网络-Transformer给出了最强的整体分类性能(准确率=0.8266),而对于这个低维表格问题,提升树仍然具有很强的竞争力。在恒星形成与AGN的二元比较中,卷积神经网络-Transformer实现了1类与4类的ROC AUC为0.9998。缺失指标提供了大量预测信息,尤其是OH_P50N_missing特征。特征解释进一步表明,log([Ne III]/[O II])与恒星质量或特定恒星形成率相结合,将恒星形成星系与AGN宿主区分开来。这些模型作为AGN候选体排名工具最有用,可补充而非取代传统的BPT诊断。高排名样本可达到高纯度,而更广泛的候选列表可找回大多数AGN,但转移到其他巡天需要进一步验证。
英文摘要
The Baldwin--Phillips--Terlevich (BPT) diagram is widely used to classify galaxies into star-forming systems, composite galaxies, and active galactic nuclei (AGNs), but its survey-scale application is limited by the requirement for high signal-to-noise emission-line measurements. We test whether machine-learning models can reproduce six-class BPT labels while using measured quantities, derived line ratios, and potentially informative missing-data patterns as inputs. We analyze 1.47 million galaxies with a 27-dimensional feature set that combines raw survey measurements, derived quantities, and missingness indicators. Five deep tabular architectures are benchmarked against gradient-boosted trees and classical machine-learning baselines, and the resulting probabilities are evaluated through hard-classification metrics, precision--recall curves, top-$k$ retrieval, ablation tests, and feature-interpretation diagnostics. {CNN--Transformer gives the strongest overall classification performance (accuracy = 0.8266), while boosted trees remain highly competitive for this low-dimensional tabular problem. In the binary star-forming-versus-AGN comparison, CNN--Transformer achieves a Class~1 versus Class~4 ROC AUC of 0.9998. Missingness indicators provide substantial predictive information, especially the OH\_P50N\_missing feature. Feature interpretation further shows that $\log([\mathrm{Ne\,III}]/[\mathrm{O\,II}])$, combined with stellar mass or specific star-formation rate, separates star-forming galaxies from AGN hosts. The models are most useful as AGN candidate-ranking tools that complement, rather than replace, traditional BPT diagnostics. High-ranked samples can reach high purity, while broader candidate lists recover most AGNs, but transferability to other surveys requires further validation.