AI 中文总结
本研究提出“冒名者”预文本任务,利用跨实体特征置换学习物理一致性,经ERA5-Land数据评估,该任务与现有SSL目标结合可提供互补信息,为科学基础模型提供新自监督来源。
AI 中文摘要
科学数据通常描述其特征受物理定律共同支配的实体,但现有的自监督学习(SSL)目标大多忽略了这种物理一致性。我们提出了“冒名者(imposter)”这一判别式预文本任务,该任务将一个实体的部分特征替换为另一个实体提供的真实观测值,并训练编码器识别被置换的特征。由于每个提供的值单独来看都是合理的,因此只有学习跨特征的物理依赖关系才能解决该任务。我们使用包含21个环境变量的全球ERA5-Land再分析数据对所提出的目标进行评估,并在涵盖气候分类、碳通量估算和径流预测的7项下游任务上评估学习到的表征。据我们所知,本研究首次在共享架构和预训练预算下,对用于陆面建模的自监督目标进行了系统比较。我们发现,最有效的预文本任务取决于下游任务类别,而非任何单一目标的优越性,且“冒名者”与现有SSL目标结合时可提供互补信息。这些结果表明,物理一致性是科学基础模型的一种宝贵的新自监督来源。
英文摘要
Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. We introduce imposter, a discriminative pretext task that replaces subsets of an entity's features with real observations donated by another entity and trains the encoder to identify the swapped features. Because every donated value is individually plausible, the task can only be solved by learning cross-feature physical dependencies. We evaluate the proposed objectives on global ERA5-Land reanalysis data using 21 environmental variables and assess the learned representations on seven downstream tasks spanning climate classification, carbon flux estimation, and streamflow prediction. Our study includes, to our knowledge, the first systematic comparison of self-supervised objectives for land-surface modeling under a shared architecture and pre-training budget. We find that the most effective pretext task depends on the downstream task family rather than any single objective's superiority, and that imposter provides complementary information when combined with existing SSL objectives. These results suggest that physical coherence is a valuable new source of self-supervision for scientific foundation models.