发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对城市理解中跨区域泛化与语义可解释性不足的问题,提出基于对比学习的时空框架CoST,通过空间相关性建模与多时相语义对齐提升性能,在8个城市指标任务中平均相对提升8.7%。
AI 中文摘要
从卫星图像学习地理空间表征是大规模城市分析及实际应用的基础问题。尽管近期已有进展,但现有方法因依赖区域特定辅助数据,且忽略多时相城市图像内的语义对齐,在跨区域泛化和语义可解释性方面表现不佳。为此,我们提出CoST,一种新颖的基于对比学习的时空框架,该框架将空间上下文与多时相语义对齐,以提取跨区域共享的通用地理规律。具体而言,CoST显式建模空间相关性以捕捉可迁移的地理结构,并利用多年城市变化语义将学习到的表征与高级地理语义对齐。大量实验表明,CoST在各类下游任务及未见过的场景中均持续取得优异性能,在8个城市指标设置下,相比最强的竞争方法平均相对提升达8.7%。代码可在该仓库获取。
英文摘要
Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel \underline{Co}ntrastive-based \underline{S}patial-\underline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7\% over the strongest competing methods across eight city-indicator settings. The code is available in \href{https://github.com/Arandinglv/CoST}{this repo}.