arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从文本中查找卫星档案中的变化:如何高效结合前后图像

Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

Simon Roy, Mark Bong, Giovanni Beltrame

arXiv 2607.28571首次发表:更新:

发表机构

Polytechnique Montréal(蒙特利尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对卫星图像变化文本查询的融合模块,对比注意力、Mamba、TBF等设计,提出两阶段搜索方案可降查询成本,TBF可降参数量与延迟,明确了不同方法的性能与效率特点。

AI 中文摘要

业务地球观测日益需要回答诸如“找到出现新建筑的图像对”之类的查询,这意味着要搜索前后时相(双时相)卫星图像对档案,并根据每对与自然语言变化描述的匹配程度进行排名。执行此匹配的组件(即结合“前”和“后”视图的融合模块)必须在查询时针对许多候选对运行,因此其速度在很大程度上决定了每次搜索的成本。我们对该模块的构建方式进行了受控比较,使用一个固定的图像编码器(冻结的CLIP模型)和所有变体的统一训练方案,评估了三种类别中的八种设计:注意力、状态空间模型(Mamba)和学习型压缩(我们的时间瓶颈融合TBF)。每种设计在两个基准(LEVIR-CC和Dubai-CC)上用十个随机种子进行测试,因此报告的差异具有统计依据。我们总结了三项发现:第一,无需训练的两阶段搜索(廉价的差异模型筛选候选,再通过注意力融合重新排名)在LEVIR-CC上的召回率与全融合相当或更高,同时将查询成本降低了10-15倍,在Dubai-CC上的R@1/R@5指标相当;第二,理论上具有吸引力的Mamba线性时间扫描,在视觉Transformer典型的补丁数量(L=196)下没有速度优势,该扫描受内存带宽限制,而注意力能很好地适配并行硬件;第三,压缩融合表示(TBF)使参数减少2.3倍,延迟降低1.6倍,变化相关的BLEU-1代价为0.007,不过更激进的压缩会悄悄丢弃与变化相关的细节,而聚合指标无法揭示这些细节。

英文摘要

Operational Earth observation increasingly calls for answering queries such as ``find the image pairs where a new building appeared.'' This means searching an archive of before-and-after (bi-temporal) satellite image pairs and ranking each pair by how well it matches a natural-language description of the change. The component that performs this match, the fusion module that combines the ``before'' and ``after'' views, must be run at query time across many candidate pairs, so its speed largely sets the cost of every search. We present a controlled comparison of how to build that module. Using one fixed image encoder (a frozen CLIP model) and one training recipe for all variants, we evaluate eight designs drawn from three families: attention, state-space models (Mamba), and learned compression (our Temporal Bottleneck Fusion, TBF). Each design is tested on two benchmarks (LEVIR-CC and Dubai-CC) with ten random seeds, so the reported differences are statistically grounded. We outline three findings: first, a training-free two-stage search (a cheap difference model that shortlists candidates, followed by attention fusion that re-ranks them) matches or exceeds full-fusion recall on LEVIR-CC while cutting query cost $10$-$15\times$, with comparable R@1/R@5 on Dubai-CC; second, the linear-time scan of Mamba, attractive on paper, gives no speed benefit at the patch counts typical of vision transformers ($L{=}196$): the scan is limited by memory bandwidth, whereas attention maps cleanly onto parallel hardware; and third, compressing the fused representation (TBF) reduces parameters by $2.3\times$ and latency by $1.6\times$ for a change-only BLEU-1 cost of $0.007$, although more aggressive compression quietly discards change-relevant detail that aggregate metrics fail to reveal.

Comments10 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑