arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

限制而非重新训练:用于零样本航拍图像分割的推理时间视觉语言模型(VLM)引导

Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation

Teresa DiMeola, Charles Walter, Hong Xiao

arXiv 2609.00628首次发表:更新:

发表机构

University of Mississippi(密西西比大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对预训练基础模型用于零样本航拍图像分割时遗漏类别与小型物体的问题,提出推理时间VLM引导方法,融合冻结基础模型与VLM查询,在四个航拍数据集上实现一致性能提升。

AI 中文摘要

全球福祉常依赖航拍与卫星图像的正确解读,基于此类图像开展的行动(如绘制淹没区域、作物范围或受损基础设施)需像素级分割以确保类别的精准定位。直接应用预训练通用基础模型时,其常遗漏重要特征,且无法总能识别给定场景的所有类别,尤其会忽略最关键的小型物体。我们采用一块消费级GPU运行视觉语言模型(VLM)以提供缺失的引导,在提升分割性能的同时生成结构化、可审计的证据,该证据可驱动结果并独立接受检查。我们融合三种方法:用于标注所有像素的冻结基础模型,以及向VLM发起的两个查询——一个用于选择重要类别,另一个用于定位基础模型遗漏的小型物体。在四个航拍数据集上的评估显示,当基础模型具备相应能力时,各阶段均取得了一致的性能提升。

英文摘要

Global welfare often depends on the correct interpretation of aerial and satellite imagery. Acting on such imagery (mapping flooded ground, crop extent, or damaged infrastructure) demands pixel-level segmentation to ensure perfect class localization. Pretrained general foundation models, when applied directly, often miss important features and cannot always find all the classes belonging to a given scene, overlooking smaller objects that matter most. We use a single consumer-grade GPU running a vision-language model (VLM) to supply this missing guidance, improving segmentation while producing structured, auditable evidence that drives the result and can be inspected on its own. We fuse three approaches: the frozen foundation model that labels every pixel, and two queries to a VLM, one to choose the classes that matter, and one to locate the small objects the base model misses. Evaluating across four aerial datasets, we see consistent gains at each stage where the base model is competent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑