JW-SSD:用于细粒度太阳黑子分类的多模态基准数据集
JW-SSD: A Multimodal Benchmark Dataset for Fine-Grained Sunspot Classification
浏览论文内容
中文总结 AI 辅助
该研究提出用于太阳黑子细粒度磁型分类的多模态基准数据集JW-SSD,经质量控制后含36553对图像,可用于训练多种模型,包括取得高准确率的JW-SunSpot。
中文摘要 AI 辅助
准确的太阳黑子分类对于评估太阳活动区的爆发潜力和预报空间天气至关重要。我们提出JW-SSD,这是一个用于太阳黑子细粒度磁型分类的高质量多模态基准数据集,由SDO/HMI SHARP 720s数据(2010-2023年,第24和25太阳活动周)构建而成,包含来自2507个活动区的36553对配准后的磁图-连续谱图像。与传统的三类分类方案不同,JW-SSD将威尔逊山分类细化为5个具有物理意义的类别(α、β、β-δ、β-γ、β-γ-δ),能够更精细地表征磁场复杂性。严格的质量控制措施包括中央子午线距离限制、饱和过滤和清晰度筛选,确保了数据的高有效性。该数据集提供FITS和PNG两种格式,包含标准的训练集(29243个)和测试集(7310个)。使用四种代表性架构(U-Net、ResNet-50、EfficientNet-B0和ViT-Small)进行的基准实验在所有模型上均取得了高准确率(三类分类任务上为89.43%-94.78%),证实该数据集在不同建模范式下均具备可靠的可学习性。JW-SSD还被用于训练多模态大语言模型JW-SunSpot,该模型实现了最高的分类准确率,证明了该数据集对传统网络和基于大语言模型的方法均具有广泛适用性。
英文摘要
Accurate sunspot classification is essential for assessing the eruptive potential of solar active regions and forecasting space weather. We present JW-SSD, a high-quality multimodal benchmark dataset for fine-grained magnetic-type classification of sunspots. Constructed from SDO/HMI SHARP 720s data (2010-2023, Solar Cycles 24 and 25), JW-SSD comprises 36,553 co-registered magnetogram-continuum pairs from 2,507 active regions. Unlike conventional three-class schemes, JW-SSD refines the Mount Wilson classification into five physically meaningful categories (α, \b{eta}, \b{eta}-δ, \b{eta}-γ, \b{eta}-γ-δ), enabling finer characterization of magnetic complexity. Rigorous quality control-including central meridian distance restriction, saturation filtering, and sharpness screening-ensures high data validity. The dataset is provided in both FITS and PNG formats, with standard training (29,243) and test (7,310) splits. Benchmark experiments with four representative architectures (U-Net, ResNet-50, EfficientNet-B0, and ViT-Small) yield high accuracy across all models (89.43%-94.78% on the three-class task), confirming that the dataset is reliably learnable across diverse modeling paradigms. JW-SSD has further been employed to train JW-SunSpot, a multimodal large language model that achieves the highest classification accuracy, demonstrating the dataset's broad applicability to both conventional networks and large-language-model-based approaches.