SLED:通过知识蒸馏实现可扩展位置编码
SLED: Scalable Location Encoding via Distillation
浏览论文内容
中文总结 AI 辅助
针对现有位置编码器计算成本高、扩展性差等问题,提出基于蒸馏的SLED,可灵活融入多模态地理空间数据,在19项基准任务上性能优于或媲美现有模型,大幅降低预训练成本。
中文摘要 AI 辅助
大量现成的地理空间数据为学习高质量的地球表征提供了绝佳机会,但地球观测(EO)数据的庞大规模、不同模态以及不同传感器类型带来了重大挑战。位置编码器已成为将EO压缩为位置特定嵌入的高效方式。然而,当前最先进的位置编码器依赖计算成本高昂的CLIP式框架,需要16K至32K的大批次规模,存在假负样本问题,且在添加其他模态时扩展性差。我们提出了通过知识蒸馏实现的可扩展位置编码器(SLED),这是一种基于蒸馏的位置编码器,利用地理空间位置作为绑定模态,可使用任意模态的地理空间数据预训练位置编码器。所得的位置编码器框架轻量、模块化,可灵活融入多种模态,同时无需样本的时空配准。SLED在批次规模小至128时仍性能良好,预训练的运行时间和计算成本仅为当前最先进模型的一小部分。我们通过在Sentinel-1、Sentinel-2和Landsat影像上预训练单模态和多模态SLED模型来验证方法,结果显示,在19项多样化的以人类为中心的基准任务上,单模态和多模态SLED模型均与现有方法相当或更优,并探究了在预训练中使用额外模态的益处。
英文摘要
The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so. Location encoders have emerged as an efficient way of compressing EOs into location-specific embeddings. However, current state-of-the-art location encoders rely on computationally expensive CLIP-style frameworks that require large batch sizes in the 16K--32K range, suffer from false negative samples, and scale poorly with additional modalities. We introduce the Scalable Location Encoder via Distillation (SLED), a distillation-based location encoder that uses geospatial location as a binding modality to pretrain location encoders with any modality of geospatial data. The resulting location encoder framework is lightweight, modular, and can flexibly incorporate multiple modes, while eliminating the need for spatiotemporal coregistration of samples. SLED is performant with batch sizes as small as 128, enabling pretraining at a fraction of the runtime and compute costs of current state-of-the-art models. We demonstrate our approach by pretraining unimodal and multimodal SLED models on Sentinel-1, Sentinel-2, and Landsat imagery. We show that both unimodal and multimodal SLED models keep pace with or outperform existing approaches on a diverse set of 19 human-centric benchmark tasks and explore the benefits of using additional modes in pretraining.