AI 中文总结
提出梯度引导解耦自适应(G2DA)框架,通过跨模态解耦和模态特定课程优化多任务地理空间视觉-语言模型,在24种基准-模型组合上显著提升性能。
AI 中文摘要
现有的地理空间视觉-语言模型(Geo-VLMs)通常通过统一的多任务自适应范式来优化多样化的地理空间任务,而没有明确考虑异构的优化特性。我们的实证观察揭示了任务间存在异构的梯度特征,包括视觉-语言差异、分支内梯度关系以及任务干扰,这些特征阻碍了有效的多任务优化。受这些观察的启发,我们提出了梯度引导的解耦自适应(G2DA),一种用于多任务Geo-VLM学习的梯度感知优化框架。G2DA首先通过梯度引导的跨模态解耦将任务划分为视觉中心和语言中心两组。然后,它基于任务梯度相似性构建模态特定的课程,并采用双向排练来缓解顺序优化引入的近因效应。我们使用六种InternVL3和Qwen3.5-VL变体以及GeoChat和GeoLLaVA在三个Geo-VLM基准上评估G2DA。在所有24种基准-模型组合中,G2DA始终优于代表性基线,在UrBench-MCQ、XLRS-Bench-Lite和VRS-Bench-VQA上分别比最强竞争对手提高了3.08、4.30和2.81个百分点。这些结果证明了梯度引导的任务组织对Geo-VLM自适应的有效性。
英文摘要
Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristics across tasks, including vision-language differences, intra-branch gradient relationships, and task interference, which hinder effective multi-task optimization. Motivated by these observations, we propose Gradient-Guided Decoupled Adaptation (G2DA), a gradient-aware optimization framework for multi-task Geo-VLM learning. G2DA first partitions tasks into vision- and language-centric groups through gradient-guided cross-modal decoupling. It then constructs modality-specific curricula based on task gradient similarity and employs bidirectional rehearsal to mitigate the recency effects introduced by sequential optimization. We evaluate G2DA on three Geo-VLM benchmarks using six InternVL3 and Qwen3.5-VL variants, along with GeoChat and GeoLLaVA. Across all 24 benchmark-model combinations, G2DA consistently outperforms representative baselines, improving over the strongest competitor by 3.08, 4.30, and 2.81 percentage points on UrBench-MCQ, XLRS-Bench-Lite, and VRS-Bench-VQA, respectively. These results demonstrate the effectiveness of gradient-guided task organization for Geo-VLM adaptation.
Comments8 pages, 3 figures, 4 tables