发表机构
University of Geneva(日内瓦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DINOspec框架通过轻量适配器和对比学习,对齐冻结的DINOv3图像编码器与预训练的AION-1光谱分词器,在20472对天体数据上提升了星系形态和光谱分类性能,验证了科学基础模型可通过轻量表征对齐组合。
AI 中文摘要
天文观测提供了物理系统的多模态视图,图像和光谱分别捕捉了天体的互补属性。科学基础模型可从这些观测中学习强大的表征,但由独立模型学习到的表征仍难以结合。我们研究能否在不重新训练视觉和光谱模型编码器的情况下,对齐这两个独立模型学习到的物理表征。我们提出了DINOspec,这是一个多模态框架,它利用轻量适配器和对比学习,在20472对天体图像与光谱上,对齐冻结的DINOv3图像编码器和预训练的AION-1光谱分词器。DINOspec在仅训练最多2100万个参数的情况下,将星系形态分类的F1值从0.72提升至0.78,将光谱分类的F1值从0.70提升至0.74。改进效果取决于下游任务,揭示了独立学习表征间的非对称迁移,而光谱红移预测保持不变(R²≈0.9)。这些结果表明,科学基础模型可通过轻量表征对齐进行组合。
英文摘要
Astronomical observations provide multimodal views of physical systems, with images and spectra capturing complementary properties of celestial objects. Scientific foundation models can learn powerful representations from these observations, but representations learned by separate models remain difficult to combine. We investigate whether physical representations learned by separate vision and spectral models can be aligned without retraining their encoders. We introduce DINOspec, a multimodal framework that aligns a frozen DINOv3 image encoder with a pre-trained AION-1 spectral tokenizer using lightweight adapters and contrastive learning on 20,472 paired images and spectra of astronomical objects. DINOspec improves galaxy morphology classification (F1: 0.72$\rightarrow$0.78) and spectral classification (F1: 0.70$\rightarrow$0.74) while training at most 21M parameters. Improvements depend on the downstream task, revealing asymmetric transfer between independently learned representations, while spectroscopic redshift prediction remains unchanged ($R^2\approx0.9$). These results demonstrate that scientific foundation models can be composed through lightweight representation alignment.
Comments4 pages, 1 figure, submitted to NeurIPS Representations for the Physical Sciences Workshop