arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22205cs.CVcs.AI

推进前填充:能力差距驱动的场景专用遥感多模态大语言模型的训练后处理

Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

Yuheng Zong, Minghua Wang, Xin Zhao, Zhi-Hui Zhan, Antonio Plaza, Jon Atli Benediktsson

首次发表
浏览论文内容

中文总结 AI 辅助

针对地球观测应用中遥感多模态大语言模型因高质量数据稀缺和能力覆盖不完整难以实现细粒度场景专业化的问题,提出能力差距驱动的推进前填充(FBA)方法,经实验验证其效果优于其他方法,提升了模型在相关基准上的表现。

中文摘要 AI 辅助

遥感多模态大语言模型(RS-MLLMs)提升了一般航空图像理解能力。然而,地球观测应用需要细粒度场景专业化,但受高质量场景数据稀缺和能力覆盖不完整的限制。我们将这种适应问题表述为能力差距驱动的训练后处理问题,并提出推进前填充(FBA)方法。FBA 不是依赖对目标域样本的单阶段监督微调,而是在推进场景专业化之前先填补先决能力差距。我们通过构建 CPRS(海岸港口遥感)数据集,将 FBA 应用于海岸港口理解这一代表性多源场景,CPRS 是一个三层监督数据集,包括三个有序阶段:(1)用于俯视视觉语言对齐 的 RS 语义锚定;(2)用于跨不同模态的目标和桥梁场景共享 RS 先验的域桥收敛;(3)用于下游性能的基于证据的场景调整。我们构建了 HarborEval,这是一个涵盖感知、空间理解、鲁棒性和生成的八轨诊断基准。在可比训练预算下,HarborEval 上,使用 FBA 时 LLaVA-v1.5 的得分从直接监督微调(Direct-SFT)的 57.95 提高到 70.29,Qwen3-VL 的得分从 81.09 提高到 83.37。FBA 还优于 Collapsed-SFT,并在与港口相关的 VRSBench/RSVQA 子集和 OpenEval 上领先。阶段分析和角色替换分析验证了渐进式差距填充和特定阶段的作用。可在指定网址获取 CPRS、HarborEval、代码和训练权重的公开示例及更新。

英文摘要

Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation applications require fine-grained scenario specialization, constrained by scarce high-quality scenario data and incomplete capability coverage. We formulate this adaptation as a capability-gap-driven post-training problem and propose filling before advancing (FBA). Rather than relying on single-stage supervised fine-tuning (SFT) over target-domain samples, FBA first fills prerequisite capability gaps before advancing toward scenario specialization. We instantiate FBA for coastal harbor understanding, a representative multi-source scenario, by constructing CPRS (Coastal-Port Remote Sensing), a three-layer supervision dataset coupled with three ordered stages: (1) RS semantic anchoring for overhead-view visual-language alignment; (2) domain-bridge convergence for shared RS priors across target and bridging scenarios under different modalities; and (3) evidence-grounded scenario tuning for downstream performance. We construct HarborEval, an eight-track diagnostic benchmark covering perception, spatial understanding, robustness, and generation. Under comparable training budgets, HarborEval increases from 57.95 with Direct-SFT to 70.29 with FBA on LLaVA-v1.5, and from 81.09 to 83.37 on Qwen3-VL. FBA also outperforms Collapsed-SFT and leads on harbor-related VRSBench/RSVQA subsets and OpenEval. Stage-wise and role-replacement analyses validate progressive gap filling and stage-specific roles. Public examples and release updates for CPRS, HarborEval, code, and trained weights are available at https://github.com/Z0ngL1ng/filling-before-advancing.

↑