发表机构
Space Applications Centre, ISRO; Indian Institute of Science Education and Research (IISER) Bhopal; Indian Institute of Technology Bombay(印度空间研究组织空间应用中心; 印度科学教育与研究学院(博帕尔); 印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对遥感基础模型与自然语言对齐不足的问题,提出受JEPA启发的AlignJEPA框架,采用轻量级预测对齐网络,结合语义预测与双向对比检索,实现了参数高效的视觉-语言对齐。
AI 中文摘要
遥感(RS)基础模型可提供跨传感器、分辨率和地理区域的可迁移地球观测表征,但多数模型与自然语言的对齐程度较弱,限制了自然语言档案搜索、图像-文本检索及基于问题的分析。本文提出AlignJEPA,一种受JEPA启发的遥感基础模型预测性视觉-语言对齐框架。AlignJEPA使用预训练的AnySat视觉编码器和RemoteCLIP文本编码器,仅训练轻量级预测对齐网络。该框架不再仅依赖全局图像-文本对比对齐,而是从被掩码的视觉基础模型令牌中预测遥感文本嵌入;其掩码感知多尺度预测对齐器在精细、区域和全局尺度聚合可见令牌,通过跨尺度Transformer联合建模这些令牌,并利用学习到的查询池化将所得表征投影到文本空间。训练过程结合语义预测与双向对比检索。我们在该http URL上训练并评估AlignJEPA以进行自然语言Sentinel检索,在RSICD上评估跨数据集适应性,仅将RSVQA用作闭集表征探针。AlignJEPA为将地球观测基础模型与语言对齐提供了一种参数高效的途径。
英文摘要
Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most remain weakly aligned with natural language, limiting natural-language archive search, image-text retrieval, and question-conditioned analysis. We propose AlignJEPA, a JEPA-inspired predictive vision-language alignment framework for remote sensing foundation models. AlignJEPA uses a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network. Instead of relying on global image--text contrastive alignment alone, the framework predicts remote-sensing text embeddings from masked visual foundation-model tokens. Its mask-aware multi-scale predictive aligner aggregates visible tokens at fine, regional, and global scales, jointly models them with a cross-scale Transformer, and projects the resulting representation into the text space using learned query pooling. Training combines semantic prediction with bidirectional contrastive retrieval. We train and evaluate AlignJEPA on BigEarthNet.txt for natural-language Sentinel retrieval, evaluate cross-dataset adaptation on RSICD, and use RSVQA only as a closed-set representation probe. AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language.
Comments18 pages