发表机构
Anhui University(安徽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对遥感变化描述中自回归方法的曝光偏差和保守生成问题,提出MGRL-RSCC,采用双解码与多粒度奖励强化学习,提升描述质量。
AI 中文摘要
遥感变化描述(RSCC)旨在从双时相遥感图像中生成关于地表物体变化的准确且详细的语言描述,是智能遥感解译中一项关键且具有挑战性的任务。主流的自回归训练范式面临严重的曝光偏差和训练-测试分布不匹配问题,导致生成误差累积。这些方法倾向于生成保守且模板固定的描述,而忽略了细微的场景变化细节。为解决这些挑战,本文提出了一种新颖的多粒度奖励强化学习范式,称为MGRL-RSCC。具体而言,我们首先利用CNN和分层自注意力模块从双时相遥感图像中提取并增强视觉特征。然后利用Transformer解码器完成视觉到语言的转换。与现有方法不同,我们设计了一种双解码策略和两阶段联合优化方案,该方案结合了通过贪婪解码的令牌级监督学习和通过采样解码的多粒度奖励驱动的自批判强化学习。我们进一步构建了三个互补的奖励函数,涵盖语言流畅性、变化状态一致性和结构-语义相关性,以全面优化描述质量并减轻虚假和缺失的变化描述。在多个公开的RSCC基准数据集上的大量实验表明,所提出的MGRL-RSCC有效缓解了传统自回归方法中的曝光偏差和保守生成问题。源代码和预训练模型将在该https URL上发布。
英文摘要
Remote Sensing Change Captioning (RSCC), which aims to generate accurate and detailed linguistic descriptions of ground object variations from bi-temporal remote sensing images, is a critical and challenging task in intelligent remote sensing interpretation. The mainstream autoregressive training paradigm faces severe exposure bias and train-test distribution mismatch, resulting in cumulative generation errors. They tend to produce conservative and template-fixed captions while ignoring subtle scene change details. To address these challenges, this paper proposes a novel multi-granularity reward reinforcement learning paradigm, termed MGRL-RSCC. Specifically, we first leverage a CNN and hierarchical self-attention module to extract and enhance visual features from bi-temporal remote sensing images. A Transformer decoder is then utilized to complete visual-to-linguistic translation. Different from existing methods, we design a dual-decoding strategy and a two-stage joint optimization scheme, which combines token-level supervised learning via greedy decoding and multi-granularity reward-driven self-critical reinforcement learning via sampling decoding. We further construct three complementary reward functions covering linguistic fluency, change state consistency, and structural-semantic relevance to comprehensively optimize caption quality and alleviate false and missing change descriptions. Extensive experiments on multiple public RSCC benchmark datasets demonstrate that the proposed MGRL-RSCC effectively mitigates exposure bias and conservative generation problems in traditional autoregressive methods. The source code and pre-trained models will be released on https://github.com/Event-AHU/MGRL-RSCC