使 Euclid VIS 成像具备 AI 就绪性:用于形态学、异常检测和相似性搜索的可扩展流水线
Making Euclid VIS Imaging AI-Ready: A Scalable Pipeline for Morphology, Anomaly Detection, and Similarity Search
浏览论文内容
中文总结 AI 辅助
针对 Euclid 海量星系图像,提出可扩展流水线,用 DINOv2 生成标准化嵌入,支持形态学、异常检测与相似性搜索,验证了多任务下游分析的可行性。
中文摘要 AI 辅助
Euclid 任务正在提供前所未有的高分辨率星系成像数据,这给跨形态学分析的标准化和可重复复用带来了新的挑战。我们提出了一种可扩展的流水线,该流水线将已发布的 VIS 切块转换为标准的 $224\ imes224$ 模型输入,并使用预训练的 DINOv2 ViT-S/14 特征提取器,从 Galaxy Zoo Euclid (Q1) 星表中构建了一个包含 365,513 行、384 维的切块与嵌入乘积。365,513 个已发布切块的初始生产耗时约 48 小时,对应的端到端速率估计为 $2.12$ 切块/秒。该服务支持批量任务管理,而对生成产品的持久复用避免了为匹配请求重复进行马赛克提取;这些机制与发布规模的生产共同定义了本文所考虑的操作可扩展性。该表示通过留出回归、少标签分类、异常候选优先级排序和相似性检索进行评估。留出岭回归探针在四个星表量上产生 $R^2=0.377$--$0.596$。在包含验证集的总标签预算为百分之一的条件下,冻结 MLP 在三个任务上的宏 F1 为 0.802--0.861,但在强不平衡的旋涡任务上为 0.484;有限的最终块微调在三个任务上给出 0.810--0.850,在旋涡任务上为 0.520。历史异常工作流识别出 1,681 个依赖配置的候选,展示的示例说明了图像质量失败,但未估计其普遍性。这些结果表明,标准化的图像接口和可复用的预训练嵌入可以在已发布的 Euclid Q1 样本规模上支持多种下游分析,而定量性能和解释在声明的输入和评估契约下仍依赖于任务。
英文摘要
The Euclid mission is delivering an unprecedented volume of high-resolution galaxy imaging, posing new challenges for standardized and reproducible reuse across morphological analyses. We present a scalable pipeline that converts released VIS cutouts into standardized $224\times224$ model inputs and uses a pretrained DINOv2 ViT-S/14 feature extractor to construct a 365,513-row, 384-dimensional cutout-and-embedding product from the Galaxy Zoo Euclid (Q1) catalogue. Initial production of the 365,513 released cutouts required approximately 48 hours, corresponding to an estimated end-to-end rate of $2.12$ cutouts s$^{-1}$. The service supports batch task management, while persistent reuse of generated products avoids repeating mosaic extraction for matching requests; together with released-scale production, these mechanisms define the operational scalability considered here. The representation is evaluated with held-out regression, few-label classification, anomaly-candidate prioritization, and similarity retrieval. Held-out ridge probes yield $R^2=0.377$--$0.596$ across four catalogue quantities. Under a one-percent total labelled budget that includes validation, frozen MLP macro-F1 is 0.802--0.861 for three tasks but 0.484 for the strongly imbalanced spiral task; limited final-block fine-tuning gives 0.810--0.850 for those three tasks and 0.520 for spiral. The historical anomaly workflow identifies 1,681 configuration-dependent candidates, and the displayed examples illustrate image-quality failures without estimating their prevalence. These results show that a standardized image interface and a reusable pretrained embedding can support several downstream analyses at the scale of the released Euclid Q1 sample, while the quantitative performance and interpretation remain task dependent under the declared input and evaluation contracts.
发表机构
- National Astronomical Observatories, Chinese Academy of Sciences(中国科学院国家天文台)
- University of Chinese Academy of Sciences(中国科学院大学)
- National Astronomical Data Center(国家天文数据中心)
机构由 AI 辅助整理,请以论文原文为准。