HIMEC:面向遥感图像变化描述的方向变化表示与固定接口解码
HIMEC: Directional Change Representation and Fixed-Interface Decoding for Remote Sensing Image Change Captioning
查看机构详情
- University of Electronic Science and Technology of China(电子科技大学)
- University of Bonn(波恩大学)
- Bonn-Aachen International Center for Information Technology (B-IT)(波恩-亚琛信息技术国际中心(B-IT))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出HIMEC模型,通过方向变化表示与固定接口解码改进遥感图像变化描述任务,在LEVIR-CC、SECOND-CC数据集上取得优于基线的CIDEr分数,验证了方法的有效性。
中文摘要 AI 辅助
遥感图像变化描述(RSICC)将双时相图像转换为描述语义变化的句子。大多数RSICC方法直接将字幕解码器建立在融合的视觉特征上,较少研究中间变化结构和解码器接口一致性。本文提出HIMEC,结合方向变化表示(DCR)与固定接口解码。DCR将带符号差异分为外观导向、消失导向和共享上下文流后再融合;学习查询编码器将融合表示转换为视觉条件变化查询标记,构成场景解码器唯一依赖样本的记忆;仅用于训练的辅助短语解码器提供来自字幕的监督。场景解码器在训练和推理时均使用固定零输入,保持相同接口。此外,评估了一种从局部到场景的级联,该级联在训练时以教师强制的局部状态为条件,推理时以自回归状态为条件。在变化的LEVIR-CC验证对上,这些状态的平均余弦距离为0.69。匹配机制的条件设置可恢复大部分相关缺陷,而打乱状态对应关系则不会产生可检测的惩罚,这些发现仅限于所评估的级联。在三次随机种子匹配的比较中,HIMEC在LEVIR-CC上达到基于共识的图像描述评估(CIDEr)分数142.81±0.60,而直接融合特征记忆的方法为139.51±3.40;在SECOND-CC上,固定零输入和匹配机制的诊断条件分别达到75.67和76.99的CIDEr分数,而不匹配的级联为60.77。源代码将在论文发表后在此公开提供。
英文摘要
Remote sensing image change captioning (RSICC) converts bitemporal imagery into a sentence describing semantic changes. Most RSICC methods condition caption decoders directly on fused visual features, leaving intermediate change structure and decoder-interface consistency less studied. We present HIMEC, combining Directional Change Representation (DCR) with fixed-interface decoding. DCR separates signed differences into appearance-oriented, disappearance-oriented, and shared-context streams before fusion. A learned-query encoder converts the fused representation into visually conditioned change-query tokens that form the scene decoder's only sample-dependent memory. A training-only auxiliary phrase decoder supplies caption-derived supervision. With a fixed zero input, the scene decoder maintains the same interface during training and inference. Separately, we evaluate a local-to-scene cascade conditioned on teacher-forced local states during training and autoregressive states at inference. On changed LEVIR-CC validation pairs, these states have a mean cosine distance of 0.69. Regime-matched conditioning recovers most of the associated deficit, whereas permuting state correspondence causes no detectable penalty. These findings are limited to the evaluated cascade. In a matched three-seed comparison, HIMEC reaches a Consensus-based Image Description Evaluation (CIDEr) score of $142.81\pm0.60$ on LEVIR-CC, versus $139.51\pm3.40$ for direct fused-feature memory. On SECOND-CC, fixed-zero and regime-matched diagnostic conditioning reach 75.67 and 76.99 CIDEr, respectively, versus 60.77 for the mismatched cascade. The source code will be made publicly available at https://github.com/ayshaashra/HIMEC upon publication.