RS-RIE-Bench:推理引导的遥感图像编辑基准测试
RS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image Editing
查看机构详情
- Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对遥感图像编辑,引入RS-RIE-Bench基准,将任务分类,通过特定评估协议和方法评估模型,经实验发现当前模型在推理引导的遥感编辑中有局限,为未来模型提供了基准和研究方向。
中文摘要 AI 辅助
遥感图像编辑旨在根据自然语言指令修改遥感图像,同时保留地理规则和传感器观测特征。现有基准主要针对自然图像或一般视觉场景,无法完全捕捉遥感编辑所需的推理、区域控制和传感器一致性能力。为此,我们引入了RS-RIE-Bench,首个用于推理引导的遥感图像编辑基准。它将任务分为时间推理、因果推理和空间推理三类,评估协议涵盖目标区域合理性、非目标区域保留和图像质量一致性三个维度。通过跨评判一致性分析和分层专家评审证明了基于MLLM评估的可行性。对八个开源和闭源图像编辑模型的系统评估表明,当前模型在推理引导的遥感编辑中仍有明显局限性。这些结果表明RS-RIE-Bench能有效揭示当前模型在地理推理、区域控制和传感器一致生成方面的局限性,为未来遥感智能编辑模型提供了标准化基准和明确研究方向。
英文摘要
Remote sensing image editing aims to modify remote sensing images according to natural language instructions while preserving geographic rules and sensor observation characteristics. Existing benchmarks mainly target natural images or general visual scenes, and thus may not fully capture the reasoning, regional control, and sensor-consistency abilities required in remote sensing editing. To fill this gap, we introduce RS-RIE-Bench, the first benchmark for reasoning-guided remote sensing image editing. RS-RIE-Bench organizes tasks into three categories: temporal reasoning, causal reasoning, and spatial reasoning. These categories capture temporal evolution, causal consequence, and spatial imaging consistency in remote sensing scenes. The evaluation protocol covers three dimensions: target region plausibility, non-target region preservation, and image quality consistency. We further demonstrate the feasibility of MLLM-based evaluation through cross-judge consistency analysis and stratified expert review. Systematic evaluation on eight open-source and closed-source image editing models shows that current models still have clear limitations in reasoning-guided remote sensing editing. Even the strongest model achieves only 24.28\% overall accuracy under the strict joint-satisfaction criterion, while the mean relaxed joint-4 success rate across all eight models is 32.23\%. Causal reasoning and spatial reasoning remain especially challenging, and several open-source models are close to zero in some categories. These results show that RS-RIE-Bench can effectively reveal the limitations of current models in geographic reasoning, regional control, and sensor-consistent generation. It also provides a standardized benchmark and a clear research direction for future remote sensing intelligent editing models.