arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeepDISC在LSST数据预览1上:ECDFS和EDFS天区的场景级检测、去混叠及恒星/星系分类

DeepDISC on LSST Data Preview 1: Scene-Level Detection, Deblending, and Star/Galaxy Classification in the ECDFS and EDFS Fields

Prince Yadav, Yaswant Sai Ejjagiri, Xin Liu, Grant Merz, the LSST Dark Energy Science Collaboration

arXiv 2610.10986首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; National Center for Supercomputing Applications, University of Illinois Urbana-Champaign; Center for Artificial Intelligence Innovation, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校; 美国国家超级计算应用中心; 人工智能创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将场景级深度学习框架DeepDISC首次应用于真实LSST数据,联合完成源检测、去混叠与恒星/星系分类,经微调后提升了源恢复率,为DeepDISC在LSST数据上建立了基准。

AI 中文摘要

我们首次将场景级深度学习框架DeepDISC应用于真实的LSST数据,使用了扩展钱德拉南天深场(ECDFS)和欧几里得南天深场(EDFS)的Data Preview 1(DP1)叠加图像。DeepDISC在单次处理中联合完成源的检测、分割,并将其分类为恒星或星系,而非像传统流程那样将这些任务视为独立步骤。我们从在LSST模拟数据上预训练的模型开始进行热启动,使用DP1六波段图像对其进行微调,将DP1星表的延展度标记(点状对应恒星、延展对应星系)作为临时分类标签。随后,仅使用经过光谱、Gaia和测光认证的恒星与星系的精选样本,对网络的分类部分单独进行微调,最终输出恒星/星系分类结果。在严格的类别感知匹配下,针对未微调模型,对真实DP1观测的微调显著提升了源的恢复率,星系和恒星的完整性从40.2%和13%分别提升至84.2%和49.1%;在类别无关匹配下,微调模型的对应值达到84.9%和74.3%。在两个分类器训练时均未见过的公共源样本上,DeepDISC对同一领域的已发表基于测光的随机森林分类结果的误差在约1个百分点内(97.4%对98.5%);在较暗星等下,DeepDISC标记的恒星数量多于随机森林,且在这些分歧中,独立标签大多支持随机森林。在保留区域与更深的HST/CANDELS成像交叉匹配显示,DeepDISC的未识别混源占比为14.9%(95%置信区间12.5-17.8%),与我们测得的生产DP1星表的18.0%一致或略低。本研究为DeepDISC在真实LSST数据上建立了基准,并为自监督预训练及LSST数据预览2奠定了基础。

英文摘要

We present the first application of DeepDISC, a scene-level deep-learning framework, to real LSST data, using the Data Preview 1 (DP1) coadded images of the Extended Chandra Deep Field-South and the Euclid Deep Field-South fields. In a single pass, DeepDISC jointly detects, segments, and classifies sources as stars or galaxies, rather than treating these tasks as separate steps as in traditional pipelines. We warm-start from a model pretrained on LSST simulations and fine-tune it on DP1 six-band images, using the DP1 catalog's extendedness flag (point-like versus extended) as provisional class labels. The classification part of the network is then fine-tuned alone on a curated sample of stars and galaxies with spectroscopic, Gaia, and photometric identifications, so the final output is a star/galaxy classification. Fine-tuning on real DP1 observations substantially improves source recovery relative to the un-fine-tuned model, raising galaxy and star completeness from 40.2% and 13% to 84.2% and 49.1% under strict class-aware matching; under class-agnostic matching the fine-tuned model reaches 84.9% and 74.3%. On a common sample of sources that neither classifier saw in training, DeepDISC classifies to within about one percentage point of a published photometry-based Random Forest for the same field (97.4% vs. 98.5%); at fainter magnitudes DeepDISC labels more sources as stars than the Random Forest, and independent labels favor the Random Forest in most of these disagreements. Cross-matching to deeper HST/CANDELS imaging on a held-out footprint gives an unrecognized-blend fraction for DeepDISC of 14.9% (95% CI 12.5-17.8%), consistent with or modestly below the 18.0% we measure for the production DP1 catalog. This work establishes a baseline for DeepDISC on real LSST data and a foundation for self-supervised pretraining and for LSST Data Preview 2.

Comments39 pages, 21 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑