发表机构
Wuhan University; Peking University; Shanghai Artificial Intelligence Laboratory(武汉大学; 北京大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出QUAKE-CD框架,将二值变化掩码编码为语法约束的四叉树令牌序列,结合思维链与双奖励强化学习,在QUAKE-CoT上取得78.31%累积F1,优于现有方法。
AI 中文摘要
遥感中的稠密变化检测要求视觉-语言模型(VLMs)比较双时相图像并生成精确的像素级掩码。现有的VLMs大多局限于变化描述输出,少数能生成像素级掩码的模型仍依赖外部解码器或扁平文本即掩码的序列化方式,这些方法对细小且碎片化的变化效果不佳。我们提出QUAKE-CD框架,将稠密变化预测重新定义为语法可验证的结构化生成。QUAKE-CD将二值变化掩码表示为语法约束的四叉树令牌序列,使掩码在自回归生成空间内紧凑、语法可检查且可确定性解码。我们进一步构建QUAKE-CoT,将序列与基于视觉证据的思维链轨迹配对,并通过渐进式课程学习及语法门控双奖励强化学习联合优化文本推理和空间稠密预测。在QUAKE-CoT上,QUAKE-CD达到78.31%的累积F1分数,优于基于解码器和扁平文本即掩码的VLMs,同时产生更忠实的双时相推理。
英文摘要
Dense change detection in remote sensing requires vision-language models (VLMs) to compare bi-temporal images and generate accurate pixel-level masks. Existing VLMs are largely confined to change captioning outputs, and the few that produce pixel-level masks still rely on external decoders or flat text-as-mask serialization, which are less effective for small and fragmented changes. We introduce QUAKE-CD, a framework that recasts dense change prediction as syntax-verifiable structured generation. QUAKE-CD represents binary change masks as grammar-constrained quadtree token sequences, making the masks compact, syntactically checkable, and deterministically decodable within an autoregressive generation space. We further construct QUAKE-CoT, which pairs these sequences with chain-of-thought traces grounded in visual evidence, and jointly optimizes textual reasoning and spatial dense prediction through a progressive curriculum followed by grammar-gated dual-reward RL. On QUAKE-CoT, QUAKE-CD achieves 78.31% accumulated F1, outperforming decoder-based and flat text-as-mask VLMs while producing more faithful bi-temporal reasoning.
Comments26 pages, 16 figures