arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FUSAR-R1:一种用于合成孔径雷达图像智能解释的大规模推理模型

FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images

Yi Yang, Xiaokun Zhang, Yuxuan Li, Ruyi Zhang, Xinpeng Zhou, Haipeng Wang

arXiv 2607.16819首次发表:更新:

发表机构

Fudan University; Key Laboratory for Information Science of Electromagnetic Waves (MoE)(复旦大学; 电磁波信息科学教育部重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对SAR图像智能解释问题,提出FUSAR-R1模型。通过模拟专家解释构建思维链推理数据指导指令学习,赋予基本推理能力,再用强化学习策略优化输出。该模型在多种SAR解释任务中表现优于现有模型。

AI 中文摘要

近年来,大规模视觉语言模型推动了智能遥感图像解释的范式转变,在合成孔径雷达(SAR)图像解释领域取得初步进展。但SAR图像受多种因素影响,现有模型缺乏人类专家的逐步分析等能力。本文提出FUSAR-R1模型,先通过模拟专家解释过程构建显式思维链推理数据指导指令学习,赋予基本推理能力,再引入强化学习策略优化输出。实验表明,FUSAR-R1在多种SAR解释任务中优于现有多模态大规模模型。

英文摘要

In recent years, large-scale vision-language models have been driving a paradigm shift in intelligent remote sensing image interpretation. By incorporating textual semantic information, the cognitive expression, semantic understanding, and human-computer interaction capabilities of interpretation models have been significantly improved, achieving initial progress in the field of Synthetic Aperture Radar (SAR) image interpretation. However, SAR images are affected by factors such as coherent imaging mechanisms, complex scattering characteristics, speckle noise interference, and target-background coupling, resulting in complex and variable image features with significant uncertainties and specializations. Existing SAR vision-language models do not yet possess the step-by-step analysis, logical judgment, and self-correction capabilities of human experts, making it difficult to support reliable intelligent interpretation in complex scenarios. To address this issue, this paper proposes a large-scale reasoning model, FUSAR-R1, for intelligent interpretation of SAR images. The model first constructs explicit chain-of-thought reasoning data by simulating the interpretation process of human experts and uses this data to guide instruction learning, thereby endowing the model with basic reasoning capabilities. Subsequently, a reinforcement learning strategy is introduced to optimize the model's outputs based on inference results, enabling self-correction and more reliable reasoning. Experimental results demonstrate that FUSAR-R1 consistently outperforms existing multimodal large-scale models across various SAR interpretation tasks, including target detection, target counting and classification, and land-cover category recognition.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑