arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05731cs.CV

T-JEPA:一种用于学习更好遥感表示的时间联合嵌入预测架构

T-JEPA: A Temporal Joint-Embedding Predictive Architecture for Learning Better Remote Sensing Representations

Bowen Peng, Li Liu, Yongxiang Liu, Weijie Li, Jie Zhou, Zhen Liu

首次发表
浏览论文内容

中文总结 AI 辅助

T-JEPA提出时间联合嵌入预测架构,利用EO序列稀疏时间采样作为监督,通过时间间隔条件化潜在转换学习表示,在匹配预训练条件下取得静态与时间任务领先迁移性能。

中文摘要 AI 辅助

地球观测(EO)数据提供了丰富的时间监督信息,然而现有的遥感基础模型主要通过施加预定义的成对关系或聚合整体重建上下文来利用序列观测。我们寻求进一步利用EO序列中固有的稀疏且不均匀的时间采样作为监督信号。为此,我们提出了T-JEPA,一种时间联合嵌入预测架构,该架构学习时间间隔条件化的潜在状态转换。一个共享的单帧编码器处理每次观测,而一个时间预测器从掩蔽的源潜在表示和实际经过的时间估计完整的目标潜在场。在多个时间间隔上,这些预测约束将观测状态组织成结构化的潜在轨迹。非对称元数据注入减轻了捷径学习,而跨多个时间尺度的直接监督比递归展开中间状态更有效。同时,掩蔽像素重建为保留空间细节提供了补充监督。在匹配的预训练数据和吞吐量下,T-JEPA在静态和时间任务上均取得了领先的迁移性能。分析进一步表明,T-JEPA学习到的表示具有时间间隔依赖的转换可预测性和连贯的潜在动态,同时保持了强跨周期一致性、表示多样性和语义可区分性。

英文摘要

Earth observation (EO) data provide rich temporal supervision, yet existing remote sensing foundation models mainly exploit sequential observations through imposing predefined pairwise relations or aggregating holistic reconstruction context. We seek to further exploit the sparse and nonuniform temporal sampling inherent in EO sequences as supervisory signals. To this end, we propose T-JEPA, a temporal joint-embedding predictive architecture that learns time-gap-conditioned latent transitions. A shared single-frame encoder processes each observation, while a temporal predictor estimates the complete target latent field from a masked source latent representation and the actual elapsed time. Across multiple temporal intervals, these predictive constraints organize observed states into structured latent trajectories. Asymmetric metadata injection mitigates shortcut learning, and direct supervision across multiple temporal scales proves more effective than recursively rolling out intermediate states. In parallel, masked pixel reconstruction provides complementary supervision for preserving spatial details. Under matched pre-training data and throughput, T-JEPA achieves leading transfer performance on both static and temporal tasks. Analyses further reveal that T-JEPA learns representations with time-gap-dependent transition predictability and coherent latent dynamics, while maintaining strong cross-period consistency, representation diversity, and semantic discriminability.

发表机构

  • National University of Defense Technology(国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑