arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06316astro-ph.GAcs.CVcs.LG

从众包中深度学习:星系形态分类的基础

Deep learning from the crowd Fundamentals of morphological galaxy classification

Luis Enrique Sucar, Carlos del Burgo, Jonathan Serrano-Pérez

首次发表
浏览论文内容

中文总结 AI 辅助

本研究利用Galaxy Zoo 1众包数据训练CNN进行星系形态分类,发现训练全网络、利用层级结构和集成可提升准确率,高一致性下准确率超99%。

中文摘要 AI 辅助

目的。本工作的目标是调整一个深度神经网络模型,使其能够从众包标注中训练进行星系形态分类,考虑训练方案、标注者之间的一致性以及层级结构。方法。我们使用Galaxy Zoo 1作为实验测试平台,训练了一个卷积神经网络(CNN)用于星系形态的自动分类。我们分析了以下方面对分类准确率和训练效率的影响:(i)仅训练最后一层 vs. 训练整个网络;(ii)仅使用CNN分类 vs. 考虑层级结构;(iii)比较使用不同数据量和标注者一致性水平训练的模型;(iv)分阶段训练,从一个模型向另一个模型迁移知识;(v)将多个模型组合成集成。结果。从实验中,我们得出以下结果:(i)训练网络的所有层显著提高了准确率(精确匹配提高10%),相比仅训练最后一层;(ii)训练所用数据量与标注者之间的一致性水平之间存在权衡;(iii)当训练数据量减少时,使用层级结构可以提高准确率;(iv)通过迁移学习(课程学习)分阶段训练在数据有限时产生更高的准确率;(v)集成可以提高准确率;(vi)对于最困难的情况,模型达到的准确率较低,但如果考虑层级度量,我们可以为层级中较高层得出有用的结果。当训练网络的所有层并考虑标注者之间高度一致性时,准确率超过99%。结论。从众包标注训练深度学习模型比从硬标注学习涉及额外的挑战。

英文摘要

Aims. The objective of this work is to adapt a deep neural network model to perform galaxy morphological classification trained from crowd annotations, considering the training scheme, the agreement between the annotators, and the hierarchy. Methods. We use Galaxy Zoo 1 as our experimental testbed and trained a convolutional neural network (CNN) for the automatic classification of galaxies' morphologies. We analyze the impact of the following aspects on the classification accuracy and training efficiency: (i) Training only the last layer vs. training all the network; (ii) Classification with only the CNN vs. considering the hierarchy; (iii) Comparing the models trained with different amounts of data and levels of agreement between the annotators; (iv) Training by stages, transferring knowledge from one model to another; and (v) Combining several models as an ensemble. Results From the experiments, we derive the following results: (i) Training all the layers in the network significantly improves the accuracy (10% increase in exact match), compared to training only the last layer; (ii) There is a tradeoff between the amount of data and the level of agreement between the annotators used for training; (iii) Using the hierarchy can improve accuracy when the amount of training data is reduced; (iv) Training by stages through transfer learning (curriculum learning) produces higher accuracy for limited data; (v) Ensembles can improve accuracy; (vi) Models achieve a low accuracy for the most difficult cases, but, if we consider hierarchical measures, we can derive useful results for upper levels in the hierarchy. An accuracy above 99% is achieved when training all layers of the network and considering a high agreement between the annotators. Conclusions. Training deep learning models from crowd annotations involves additional challenges than learning from hard annotations.

发表机构

  • Instituto Nacional de Astrofísica, Óptica y Electrónica(国家天体物理、光学与电子学研究所)
  • Universidad de la Laguna(拉古纳大学)
  • Universidad Autónoma de Tlaxcala(特拉斯卡拉自治大学)
  • Instituto de Astrofísica de Canarias(加那利群岛天体物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑