在千度巡天中利用机器学习集成广义均值共识减少强引力透镜搜索中的假阳性
Reducing False Positives in Strong-Lens Searches with Generalized-Mean Consensus of Machine-Learning Ensembles in the Kilo-Degree Survey
- Institute for Astrophysics, School of Physics, Zhengzhou University(郑州大学物理学院天体物理研究所)
- International College, Zhengzhou University(郑州大学国际教育学院)
- School of Physics and Astronomy, Beijing Normal University(北京师范大学天文与物理学院)
- INAF – Osservatorio Astronomico di Capodimonte(意大利国家天体物理研究所卡波迪蒙特天文台)
- Department of Physics “E. Pancini”, University Federico II(那不勒斯费德里科二世大学E. Pancini物理系)
- School of Mathematics and Physics, Xi’an Jiaotong-Liverpool University(西交利物浦大学数理学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究在KiDS DR4中结合多种机器学习分类器,通过广义均值共识集成,在保持高完备性的同时大幅降低强透镜搜索的假阳性率,减少视觉检查工作量。
AI中文摘要:
背景。在广域巡天中,主要挑战不仅仅是分类器的灵敏度,还有数量庞大的假阳性。在数百万到数十亿个星系中搜索强引力透镜会产生大量污染物,使其成为后续检查和构建具有统计意义的透镜样本的瓶颈。目标。我们旨在通过组合多个分类器来提高KiDS DR4中强透镜候选体选择的纯度。目标是保持已知候选体的高完备性,同时大幅降低非透镜的比例。方法。我们训练了卷积、基于Transformer和混合分类器,包括Li ResNet+、Swin Transformer变体、Swin-MLP和DemiLensNet。它们的概率输出在分数层面通过平均和广义均值共识进行组合。模型在模拟的KiDS类透镜图像上进行了测试,然后在嵌入非透镜样本中的真实KiDS DR4透镜候选体上进行了评估。结果。在模拟测试集上,集成相对于最佳单一模型没有优势。在混合的真实KiDS测试集上,算术平均在90%完备性下将假阳性率从0.016-0.020(两个最佳单一模型所覆盖的范围)降至七模型集成的0.011。广义均值进一步将其降至0.007。在相同的90%完备性水平下应用于完整的LRG和BG样本时,相对于最佳单一模型,广义均值将返回的候选体数量分别减少了约50%(LRG)和70%(BG)。经过视觉检查,我们获得了170个新的高质量候选体(24个A类和146个B类),以及1706个C类候选体。结论。我们的结果表明,机器学习集成的广义均值共识策略提供了一条实用途径,可在保持高恢复率的有前景强透镜候选体的同时,减少视觉检查工作量。
英文摘要:
Context. In wide-field surveys, the main challenge is not just classifier sensitivity, but the overwhelming number of false positives. Searching for strong lenses among millions to bilions of galaxies produces many contaminants, making the bottleneck for follow-up inspection and building statistically useful lens samples. Aims. We aim to improve the purity of strong-lens candidate selection in KiDS DR4 by combining several classifiers. The objective is to retain high completeness for known candidates while substantially reducing the fraction of non-lenses. Methods. We trained convolutional, Transformer-based, and hybrid classifiers, including Li ResNet+, Swin Transformer variants, Swin-MLP, and DemiLensNet. Their probabilistic outputs were combined at score level using averaging and a generalized mean consensus. The models were tested on simulated KiDS-like lens images and then evaluated on real KiDS DR4 lens candidates embedded in a non-lens sample. Results. On the simulated test set, ensembles show no advantage over the best single models. On the mixed real KiDS test set, the arithmetic mean reduces the false-positive rate at 90% completeness from 0.016-0.020 (the range spanned by the two best individual models) to 0.011 for the seven-model ensemble. The generalized mean reduces it further, to 0.007. Applied to the full LRG and BG samples at the same 90% completeness level, the generalized mean reduces returned candidates by roughly 50% for LRGs and 70% for BGs, relative to the best single model. After visual inspection, we obtain 170 new high-quality candidates (24 Class A and 146 Class B), together with 1706 Class C candidates. Conclusions. Our results demonstrate that the generalized mean consensus of an ML ensemble strategy provides a practical route to reducing the visual inspection workload while preserving a high recovery rate of promising strong-lens candidates.