arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38072cs.CV

RS-OPSD:面向超高分辨率遥感视觉问答的可靠特权在线策略自蒸馏

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对超高分辨率遥感VQA中微小证据解析难题,提出RS-OPSD框架,通过特权在线自蒸馏将视觉特权内化,无需推理时额外搜索,构建GeoEvidence-6K数据集并引入CPVP与CAD,在多个基准上平均提升4.0个百分点,2B变体超越多数8B模型。

中文摘要 AI 辅助

超高分辨率(UHR)遥感视觉问答(VQA)要求模型在极大的图像中解析微小的视觉证据。现有方法通常依赖推理时的令牌剪枝、视觉搜索或工具增强推理。我们转而研究是否可以将缩放视觉特权的益处内化到模型中。我们提出了RS-OPSD,一种用于UHR遥感VQA的可靠特权在线策略自蒸馏(OPSD)框架。为了提供带有明确问题相关证据的高质量特权信息,我们构建了GeoEvidence-6K,包含七个任务类别中的6,750个VQA样本及证据区域标注,并开发了人类反馈引导的技能细化(HF-SR)以实现可扩展标注。为解决紧凑裁剪导致的上下文丢失和不完美教师带来的冲突信号,RS-OPSD引入了上下文保留视觉特权(CPVP)和正确性对齐蒸馏(CAD)。在推理时不进行任何额外的视觉搜索或工具调用,RS-OPSD在XLRS-Bench、MME-RealWorld-RS和LRS-VQA上达到了最先进的(SOTA)性能,平均比之前同规模SOTA模型高出4.0个百分点。此外,我们的2B变体RS-OPD-Lite超越了大多数8B规模模型,同时实现了最快的实测推理速度。我们的代码、GeoEvidence-6K以及RS-OPSD和RS-OPD-Lite的模型权重均已公开。

英文摘要

Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming previous SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.

发表机构

  • Tsinghua University(清华大学)
  • Zhejiang University(浙江大学)
  • Central University of Finance and Economics(中央财经大学)
  • East China Normal University(华东师范大学)
  • Key Laboratory of Geographic Information Science(地理信息科学重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑