AI 中文总结
研究AGN和SFG分类中添加X射线通量特征致准确率降低的原因,发现天体物理学上两类星系X射线光度有重叠,机器学习中训练标签有噪声,X射线特征对模型预测有损害,结论是差异源于机器学习假象与天体物理学重叠,未来分类器应采用独立标签源或抗噪声训练方法。
AI 中文摘要
在之前的一篇论文中,我们发现向活动星系核(AGN)和恒星形成星系(SFG)的随机森林分类器中添加X射线通量特征会使分类准确率从97.51%降至89.26%,鉴于普遍认为X射线是可靠的AGN诊断依据,这一结果有悖直觉。本文从天体物理学和机器学习两个角度研究了这种差异的来源。在天体物理学方面,我们表明样本中AGN和SFG的X射线光度在标准的\(10^{42}\)尔格/秒阈值上有大量重叠,这反映了光学选择的SDSS样本的中等光度特征以及高质量X射线双星(HMXB)对SFG发射的贡献。在机器学习方面,我们认为BPT衍生的训练标签构成了实例依赖的标签噪声:标签不确定性集中在BPT分界线附近,而在那里X射线数据作为判别器最为有用。我们表明,在5折交叉验证中误分类的对象主要聚集在Kewley分界线曲线附近,与正确分类的对象相比,中位数距离为0.130 dex,而正确分类对象的中位数距离为0.748 dex(Kolmogorov - Smirnov检验结果为\(D = 0.701,\; p = 3. /times 10^{-9}\),显示出统计学上的显著差异),并且X射线特征表现出与对模型预测有实际损害一致的排列重要性。我们得出结论,这种差异与由天体物理学重叠加剧的机器学习假象一致,并且未来对光学选择样本进行操作的分类器应采用独立的标签源或抗噪声训练方法。
英文摘要
In a previous paper, we found that adding an X-ray flux feature to a Random Forest classifier of active galactic nuclei (AGN) and star-forming galaxies (SFGs) coincided with a decrease in classification accuracy from 97.51% to 89.26%, a counterintuitive result given the prevailing theory that X-rays are a reliable AGN diagnostic. This paper investigates the source of that discrepancy through both astrophysical and machine learning lenses. On the astrophysical side, we show that the X-ray luminosities of AGN and SFGs in the sample substantially overlap across the canonical $10^{42}$ erg/s threshold, reflecting the moderate luminosity characteristic of an optically-selected SDSS sample and the contribution of high-mass X-ray binaries (HMXBs) to SFG emission. On the machine learning side, we argue that the BPT-derived training labels constitute instance-dependent label noise: label uncertainty is concentrated near the BPT demarcation line, where X-ray data would be most useful as a discriminator. We show that misclassified objects in 5-fold cross-validation cluster primarily near the Kewley demarcation curve, with a median distance of 0.123 dex compared to 0.743 dex for correctly classified objects (Kolmogorov-Smirnov $D = 0.745,\; p = 1.67\times 10^{-9}$). A controlled cross-validation decomposition shows that the previously reported decrease reflects predominantly sample selection rather than the X-ray feature, which exhibits a small but significantly negative permutation importance: the model makes limited, net-detrimental use of it, though its effect on overall accuracy is negligible. We conclude that future classifiers operating on optically-selected samples should employ independent label sources or noise-robust training methods.
Comments7 pages, 4 figures. Submitted to RASTI