arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Freq-RemoteVAR:用于遥感变化检测的下一个频率自回归建模

Freq-RemoteVAR: Next-Frequency Autoregressive Modeling for Remote Sensing Change Detection

Luqi Gong, Rui Xu, Yue Chen, Chao Li, Jingqi Hong, Xuefeng Zhao

arXiv 2607.25815首次发表:更新:

发表机构

Zhejiang Lab; Beijing University of Posts and Telecommunications; Changsha University of Science and Technology; The Education University of Hong Kong; South China Agricultural University(之江实验室; 北京邮电大学; 长沙理工大学; 香港教育大学; 华南农业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对遥感变化检测,提出Freq-RemoteVAR框架,将其转为频域结构化生成问题,采用下一个频率预测范式,设计频率感知掩码令牌化等策略,经实验验证该方法在复杂场景中性能优于现有方法。

AI 中文摘要

遥感变化检测旨在从双时相图像中识别土地覆盖变化。大多数现有方法采用一次性密集预测范式,直接从融合特征回归变化掩码,但忽略了变化模式的内在频率特征。我们提出了Freq-RemoteVAR,一种频率自回归框架,将变化检测重新表述为频域中的结构化生成问题。引入下一个频率预测范式,通过傅里叶变换和量化将变化监督分解为多频率令牌目标,设计频率感知掩码令牌化策略。开发频率VAR Transformer对频率令牌进行因果自回归建模,引入Scale-Aligned RoPE Cross Attention (SRCA) 模块增强空间频率一致性,提出变化质量控制模块抑制伪变化响应并提高鲁棒性。在CDD、GZ-CD和LEVIR-CD上的大量实验表明,Freq-RemoteVAR始终优于现有方法,尤其是在具有复杂外观变化和噪声干扰的具有挑战性的场景中。

英文摘要

Remote sensing change detection aims to identify land-cover changes from bi-temporal images. Most existing methods follow a one-shot dense prediction paradigm, directly regressing a change mask from fused features. However, such approaches overlook the intrinsic frequency characteristics of change patterns. We propose Freq-RemoteVAR, a frequency autoregressive framework that reformulates change detection as a structured generation problem in the frequency domain. Instead of predicting the change mask in a single step, we introduce a next-frequency prediction paradigm, where change information is progressively generated from coarse to fine. We design a frequency-aware mask tokenization strategy that decomposes change supervision into multi-frequency token targets via Fourier transformation and quantization. We develop a Frequency VAR Transformer, which performs causal autoregressive modeling over frequency tokens. The model starts from learned mask queries and progressively predicts frequency-level tokens conditioned on previously generated tokens and bi-temporal image features, effectively capturing long-range dependencies across frequency scales. We introduce Scale-Aligned RoPE Cross Attention (SRCA) module, which aligns frequency-domain mask queries with spatial-domain bi-temporal features under a unified coordinate system, enhancing spatial-frequency consistency during generation. We propose a Change-quality Control module that adaptively modulates the generation process through dynamic normalization, attention biasing, and spatial offset adjustment, thereby suppressing pseudo-change responses and improving robustness. Extensive experiments on CDD, GZ-CD, and LEVIR-CD demonstrate that Freq-RemoteVAR consistently outperforms existing methods, particularly in challenging scenarios with complex appearance variations and noisy disturbances.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑