arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15942cs.CVcs.LG

以少胜多:一种简单方法构建的大规模遥感视觉语言模型

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对遥感视觉语言模型,质疑架构专业化必要性,提出通用模型经大规模跨数据和任务训练可获好性能。核心方法是用单一语言策略及多任务强化学习框架训练。主要贡献是在多基准测试中取得竞争力结果,证明数据规模更关键。

中文摘要 AI 辅助

遥感视觉语言模型越来越被期望支持对地球观测数据和各种任务进行开放式推理。该领域的最新进展往往依赖于特定于遥感的架构设计。本文对这种架构专业化的必要性提出质疑,表明一个通用的视觉语言模型在经过足够规模的跨多样数据和任务训练后,能在具有挑战性的遥感基准测试中取得有竞争力或领先的性能。模型采用单一语言策略,可直接文本回答或调用定位工具进行分割和定位。通过多任务强化学习框架及自适应任务奖励进行训练,涵盖多种输入类型的多项任务。该方法在众多基准测试中取得了有竞争力的结果,且随着训练数据规模增加,多数任务表现持续提升,表明数据规模对遥感视觉语言模型比架构新颖性更重要。

英文摘要

Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is sufficient even without architectural novelty.

发表机构

  • INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学“圣克莱门特·奥赫里德斯基”信息与自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑