发表机构
Georgia Institute of Technology; Indian Institute of Engineering Science and Technology (IIEST), Shibpur; Nanyang Technological University; Indian Institute of Technology, Delhi; Duke University(佐治亚理工学院; 印度工程科学与技术学院(IIEST),希布普尔; 南洋理工大学; 印度理工学院德里分校; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BMD-CD,结合时序状态空间建模与logit空间扩散细化,实现高效高精度的遥感变化检测,在多个基准上取得领先性能。
AI 中文摘要
遥感变化检测需要对双时相图像进行全局推理,并精确定位变化区域。然而,对于高分辨率图像,密集注意力计算成本高昂,而传统的特征融合和粗解码可能无法充分区分真实变化与外观变化,或难以保留物体边界。我们提出了用于变化检测的双时相Mamba-扩散模型(BMD-CD),该模型将时序结构化的状态空间建模与logit空间扩散细化相结合。BMD-CD将深层双时相特征转换为区域令牌,并在双向状态空间传播之前将其排列为显式的时序分区。其双时相有序Mamba算子能够以线性序列复杂度实现长距离跨时相交互,而正交特征解缠通过学习到的成对旋转和未变化区域一致性,形成面向变化的输出和互补的旋转输出。多尺度解码随后产生粗略的变化logits,并通过五步条件扩散解码器直接在logit空间中进行细化。在LEVIR-CD、WHU-CD、DSIFN-CD、CDD和S2Looking上的实验表明,该方法在多种变化检测设置下均表现出强大的性能。BMD-CD在四个标准基准上分别取得了93.7%、96.0%、97.8%和99.0%的F1分数,并在LEVIR-CD和WHU-CD上将3像素边界F1分别提升至87.7%和91.4%。完整模型每对256×256图像需要32.09 GFLOPs和47毫秒的计算量,同时也在ValaisCD和B-FLAIR-test上展示了零样本迁移能力。我们的代码可在https://github.com/Aparup2139/Public_WACV/获取。
英文摘要
Remote sensing change detection requires both global reasoning across bitemporal images and precise localization of changed regions. However, dense attention is computationally expensive for high-resolution imagery, while conventional feature fusion and coarse decoding may inadequately separate genuine changes from appearance variations or preserve object boundaries. We present Bitemporal Mamba-Diffusion for Change Detection (BMD-CD), which combines temporally structured state-space modeling with logit-space diffusion refinement. BMD-CD converts deep bitemporal features into region tokens and arranges them in explicit temporal partitions before bidirectional state-space propagation. Its Bitemporal Ordered Mamba Operator enables long-range cross-temporal interaction with linear sequence complexity, while Orthogonal Feature Disentanglement forms a change-oriented output and a complementary rotated output using learned pairwise rotations and unchanged-region consistency. Multiscale decoding then produces coarse change logits, which are refined through a five-step Conditional Diffusion Decoder operating directly in logit space. Experiments on LEVIR-CD, WHU-CD, DSIFN-CD, CDD, and S2Looking demonstrate strong performance across diverse change-detection settings. BMD-CD achieves F1 scores of 93.7%, 96.0%, 97.8%, and 99.0% on the four standard benchmarks and improves 3-pixel Boundary-F1 to 87.7% and 91.4% on LEVIR-CD and WHU-CD, respectively. The full model requires 32.09 GFLOPs and 47 ms per 256 x 256 image pair, while also showing zero-shot transfer to ValaisCD and B-FLAIR-test. Our code is available at https://github.com/Aparup2139/Public_WACV/
CommentsSubmitted to WACV2027 Application Track