arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有类别不确定性的天文光谱贝叶斯分类

Bayesian classification of astronomical spectra with class uncertainties

Simon Barton, Martin Sahlén, Andreas Korn, Christian Glaser

arXiv 2609.21694首次发表:更新:

发表机构

Uppsala University; TU Dortmund University(乌普萨拉大学; 多特蒙德工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对4MOST巡天光谱分类,比较四种贝叶斯方法,发现基于CNN的蒙特卡洛dropout在准确率、校准和效率上最优,分别达92.6%和93.9%。

AI 中文摘要

背景:我们开发了一种概率机器学习方法,旨在为即将开展的4MOST巡天项目,对恒星和河外天体的低分辨率和高分辨率光谱进行O(10)类分类。为满足巡天要求,该方法应能表达输入数据中的不确定性以及其预测中引入的不确定性。目标:我们探索了四种不同的方法:(1)卷积神经网络(CNNs),(2)狄利克雷分布,(3)蒙特卡洛dropout(MCD),(4)贝叶斯神经网络(BNNs)+变分推断(VI)。使用来自SDSS数据库的标记光谱和自定义的4MOST模拟数据集进行了训练和验证。所有方法都根据相同的指标进行比较:准确率、曲线下面积(AUC)、期望校准误差(ECE)、香农熵、负对数似然(NLL)、布里尔分数、训练时间和推理时间。方法:训练了一个架构简单、约20,000个参数的CNN,在SDSS数据上实现了91.5%的分类准确率,在4MOST模拟数据上实现了92.8%的分类准确率。直接狄利克雷预测和VI模型测试提供了类别成员概率的不确定性,但更频繁地混淆类别。发现CNN上的MCD最为合适;它将点估计准确率提升至92.6%和93.9%,同时仍提供快速训练和足够快的推理。与标准CNN相比,该方法以边际额外成本额外提供了良好校准的不确定性。

英文摘要

Context: We developed a probabilistic machine learning method with the aim of performing the O(10)-way classification of low- and high-resolution spectra of stellar and extragalactic targets for the upcoming 4MOST survey. In fulfilment of the survey requirements, this method should be able to express uncertainty in the input data as well as uncertainty introduced in its prediction. Aims: Four different methods are explored: (1) convolutional neural networks (CNNs), (2) the Dirichlet distribution, (3) Monte Carlo dropout (MCD), (4) Bayesian neural Networks (BNNs) + variational inference (VI). Training and validation was performed using labelled spectra from the SDSS database and a custom 4MOST mock dataset. All the methods were compared in terms of the same metrics: accuracy, area under the curve (AUC), expected calibration error (ECE), Shannon entropy, negative log-likelihood (NLL), Brier score, training time, and inference time. Methods: A CNN with simple architecture and about 20,000 parameters was trained to achieve classification accuracies of 91.5% on SDSS data and 92.8% on 4MOST mock data. The direct Dirichlet prediction and VI models tested provide uncertainties on class membership probabilities, but they confuse classes more often. The MCD on a CNN is found to be the most suitable; it boosts the point-estimate accuracies to 92.6% and 93.9%, while still providing fast training and sufficiently fast inference. Compared to a standard CNN, the method additionally provides well-calibrated uncertainties at marginal extra cost.

Journal refA&A, 713, A55 (2026)

DOI:10.1051/0004-6361/202556350

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑