arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重用还是重新学习?地球观测基础模型的谱视角

Reuse or Relearn? A Spectral View of Earth Observation Foundation Models

Mehmet Ozgur Turkoglu, Valerio Marsocci, Dominik J. Mühlematter, Dominik Senti, Konrad Schindler, Helge Aasen

arXiv 2609.32756首次发表:更新:

AI 中文总结

本文通过谱诊断方法研究地球观测基础模型微调是重用还是重新学习预训练表示,发现EO模型更新大且保留少,并指出应同时评估预训练表示的可重用性。

AI 中文摘要

基础模型很少被用作通用的、冻结的特征提取器;相反,它们会针对目标下游应用进行微调。这种做法在地球观测(EO)领域尤为普遍,并引发了一个仅凭下游精度无法回答的问题:微调是重用了预训练表示,还是重新学习了一个新的表示?我们通过谱诊断来研究这一问题,该诊断比较模型在适应前后的状态,量化其主导奇异子空间被保留的程度、权重更新的分布广度以及更新幅度的大小。以CLIP和DINO等自然图像模型作为参考,我们发现,在所评估的微调设置下,EO模型经历了更大、更高秩的更新,并且保留的预训练结构要少得多,因此它们的下游性能往往是通过对预训练权重结构进行实质性改变而获得的。这些诊断进一步提供了关于模型适应成本低廉程度的见解:在预训练子空间被保留的情况下,适应一小部分参数即可匹配完全微调的效果;而在预训练子空间未被保留的情况下,适应一小部分参数则可能落后于完全微调。更广泛地说,基础模型,尤其是EO基础模型,不仅应通过基准精度来评估,还应通过其预训练表示的可重用性来评估。

英文摘要

Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reuse the pretrained representation, or does it relearn a new one? We study this with spectral diagnostics that compare a model before and after adaptation, quantifying how well its dominant singular subspaces are preserved, how broadly the weight update is distributed, and how large it is. Using natural image models such as CLIP and DINO as a reference, we find that, under the evaluated fine-tuning settings, EO models undergo far larger, higher-rank updates and retain much less of their pretrained structure, so their downstream performance is often obtained with substantial changes to the pretrained weight structure. The diagnostics further provide insight into how cheaply a model can be adapted: where the pretrained subspaces are preserved, adapting a small fraction of the parameters can match full fine-tuning, and where they are not, it can fall behind. More broadly, foundation models, and EO foundation models in particular, should be assessed not only by benchmark accuracy, but also by how reusable their pretrained representation is.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑