arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

月球遥感的多模态-多分辨率基础模型

Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing

Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal, Rachel A. Slank, Geoffrey Dawson, Hiyam Debary, Michael K. Barker, Andrew Annex, Vishnu Viswanathan, Zachary Morse, Ethan I. Schaefer, Nikolaos Dionelis, Ankur Kumar, Campbell D. Watson, Manil Maskey, Rebekah I. Dawson-Rigas, Juan Bernabé-Moreno, Rahul Ramachandran, Sujit Roy

arXiv 2609.13283首次发表:更新:

发表机构

IBM Research; The University of Alabama in Huntsville; University Space Research Association (USRA); NASA Goddard Space Flight Center; SETI Institute; NASA Ames Research Center; University of Maryland, Baltimore County (UMBC); Howard University(IBM研究院; 阿拉巴马大学亨茨维尔分校; 大学空间研究协会; 美国国家航空航天局戈达德太空飞行中心; 搜寻地外文明研究所; 美国国家航空航天局艾姆斯研究中心; 马里兰大学巴尔的摩县分校; 霍华德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种月球遥感多模态基础模型,在SomBench上预训练,通过多尺度联合训练和几何上下文,在陨石坑检测、IMP分割和极地冰回归等任务上匹配或超越ImageNet基线,并展示LoRA的高效性。

AI 中文摘要

我们提出了一种用于月球遥感的多模态基础模型,该模型在SomBench上从头开始预训练。SomBench是一个按地理划分的数据集,包含近两百万个配准的瓦片束,覆盖11种模态,空间尺度为两种(1米/像素和100米/像素)。该模型改编了TerraMind掩码令牌架构,并增加了两个针对月球的扩展:采集几何信息作为显式上下文提供,米级和百米级瓦片联合训练,使得单一权重集覆盖两种分辨率。FlexiViT补丁嵌入允许在不重新训练的情况下适应不同的补丁大小,而按模态的输入支持灵活的多模态微调。定性生成实验表明,该模型学习了有意义的跨模态对应关系,包括从高程得到的地形衍生量,以及从几何得到的与光照一致的反射率。我们在四个基准上进行了评估:WAC和NAC尺度的陨石坑检测、不规则月海斑(IMP)分割以及极地冰前景回归。在所有任务中,预训练模型匹配或优于ImageNet预训练基线以及架构相同的随机初始化对照。在多模态冰前景回归中,预训练变体取得了最佳结果,而随机初始化模型优于大多数基线,这表明性能提升既来自架构也来自预训练。在WAC陨石坑检测中,标签效率显著:在50%数据上训练的预训练模型超过了在完整数据集上训练的最强ImageNet基线。在适应策略中,LoRA在陨石坑检测和IMP分割上匹配或超过全量微调,同时使用的可训练参数少得多,而全量微调在冰前景回归中表现最佳。我们发布了预训练检查点、基准数据集和微调代码,以支持可复现的月球AI研究。

英文摘要

We present a multimodal foundation model for lunar remote sensing, pretrained from scratch on SomBench, a geographically partitioned corpus of nearly two million co-registered tile bundles spanning 11 modalities at two spatial scales (1 m/pixel and 100 m/pixel). The model adapts the TerraMind masked-token architecture with two lunar-specific extensions: acquisition geometry is provided as explicit context, and meter- and hundred-meter-scale tiles are trained jointly so that a single set of weights covers both resolutions. FlexiViT patch embeddings allow adaptation to different patch sizes without retraining, while modality-wise inputs enable flexible multimodal fine-tuning. Qualitative generation experiments suggest the model learns meaningful cross-modal correspondences, including terrain derivatives from elevation and illumination-consistent reflectance from geometry. We evaluate on four benchmarks: crater detection at WAC and NAC scales, irregular mare patch (IMP) segmentation, and polar ice prospectivity regression. Across tasks, the pretrained model matches or outperforms ImageNet-pretrained baselines and an architecturally identical random-init control. On multimodal ice prospectivity regression, pretrained variants achieve the best results, while the random-init model outperforms most baselines, suggesting gains arise from both the architecture and pretraining. Label efficiency is notable for WAC crater detection, where the pretrained model trained on 50% of the data exceeds the strongest ImageNet baseline trained on the full dataset. Among adaptation strategies, LoRA matches or surpasses full fine-tuning on crater detection and IMP segmentation while using far fewer trainable parameters, whereas full fine-tuning performs best for ice prospectivity regression. We release the pretrained checkpoint, benchmark datasets, and fine-tuning code to support reproducible lunar AI research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑