BEACON:通过三模态对比学习对AlphaEarth嵌入进行行为与语义增强
BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning
浏览论文内容
中文总结 AI 辅助
本文提出三模态对比学习框架BEACON,补充地理空间基础模型AlphaEarth的语义与行为信号,在休斯顿都会区9项下游任务上较6个基准方法提升了人类相关指标预测性能,拓展了其城市分析应用。
中文摘要 AI 辅助
AlphaEarth Foundation等地理空间基础模型能生成紧凑且全局一致的地球表面表征,可有效迁移至各类下游任务。但因这类模型主要基于地球观测影像训练,其嵌入主要捕获物理与光谱特征,对人类活动及城市功能的编码能力较弱。为解决此局限,本文提出BEACON框架,这是一种三模态对比学习框架,对齐城市空间的三类互补视角:来自AE嵌入的物理表征、来自兴趣点(POI)文本的语义表征,以及来自小时级POI访问量的人类行为表征,同时保持部署的表征仅为图像形式。本文以休斯顿都会区为案例研究区域,在9项下游任务(含7项回归任务、2项分类任务)上评估BEACON框架的性能,对比6个基准方法(原始坐标、Space2Vec、SatCLIP、TESSERA、Clay及AlphaEarth),采用冻结线性探针与多层感知机(MLP)探针,设置5个随机种子。在线性探针设置下,BEACON针对肥胖患病率的相对R²较AlphaEarth提升最高达43%,针对精神健康不佳提升34%,针对家庭收入中位数提升22%,同时在物理与环境变量预测上保持竞争力。这些发现凸显了为地理空间基础模型补充语义与行为信号的价值,将其应用范围从物理地球观测扩展至以人类为中心的城市分析。
英文摘要
Geospatial foundation models such as the AlphaEarth Foundation produce compact and globally consistent representations of the Earth's surface that transfer effectively to a wide range of downstream tasks. However, because these models are trained primarily on Earth-observation imagery, their embeddings mainly capture physical and spectral characteristics while encoding human activity and urban function only weakly. To address this limitation, we propose BEACON, a tri-modal contrastive learning framework that aligns three complementary views of urban space: physical representations from AE embeddings, semantic representations from point-of-interest (POI) text, and human behavioral representations from hourly POI visitation, while keeping the deployed representation image-only. Using the Houston Metropolitan Area as a case study area, we evaluated the performance of the BEACON framework on nine downstream tasks, including seven regression and two classification tasks against six baselines (raw coordinates, Space2Vec, SatCLIP, TESSERA, Clay and AlphaEarth), using frozen linear and MLP probes over five seeds. Under a linear probe, BEACON improves relative R^2 over AlphaEarth by up to 43% for obesity prevalence, 34% for poor mental health, and 22% for median household income, while remaining competitive in the prediction of physical and environmental variables. These findings highlight the value of augmenting geospatial foundation models with semantic and behavioral signals, extending their applicability from physical Earth observation to human-centered urban analytics.
发表机构
- Texas A&M University(得克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。