arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29367cs.CV

SatEdit:基于VLM引导的区域标注的掩码条件卫星图像编辑

SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

Muhammad Talha, Muhammad Ahmed Amer

首次发表
浏览论文内容

中文总结 AI 辅助

SatEdit是基于VLM引导区域标注的卫星图像编辑框架,通过未标注图像构建监督信号,经LoRA微调后在对比中实现最高掩码区域语义对齐,为卫星图像编辑提供了数据高效的可行方案。

中文摘要 AI 辅助

卫星图像编辑需要空间精确的物体级控制,但用于高空图像的监督编辑数据集构建成本高昂,因为物体掩码、语义标签和配对编辑内容很少能大规模获取。我们提出SatEdit,一种基于掩码条件的卫星图像编辑框架,该框架从未标注图像中构建训练监督信号。SatEdit利用分割基础模型生成物体掩码,通过视觉语言模型(VLM)为采样区域分配语义标签,并在通过掩码引导的图像修复生成配对添加和移除样本前进行轻量人工验证。我们在源自SODA-A的数据集上用LoRA微调高分辨率图像编辑骨干网络,该数据集包含1014张图像和91类共852个已验证物体标注。在与开源及专有图像编辑模型的受控对比中,SatEdit实现了最高的聚合掩码区域语义对齐,其CLIP得分为0.6322,CLIP差值为0.0726,同时在视觉上保留了周围场景。这些结果表明,VLM辅助的区域标注是实现数据高效、空间可控的卫星图像编辑的可行路径。

英文摘要

Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build because object masks, semantic labels, and paired edits are rarely available at scale. We introduce SatEdit, a mask-conditioned satellite image editing framework that constructs training supervision from unlabeled imagery. SatEdit proposes object masks with a seg- mentation foundation model, assigns semantic la- bels to sampled segments with a Vision-Language Model, and applies lightweight human verification before generating paired addition and removal exam- ples through mask-guided inpainting. We fine-tune a high-resolution image editing backbone with LoRA on a SODA-A-derived dataset containing 1,014 im- ages and 852 verified object annotations across 91 classes. In controlled comparisons with open- source and proprietary image editing models, SatE- dit achieves the highest aggregate masked-region se- mantic alignment, with a CLIP score of 0.6322 and CLIP delta of 0.0726, while preserving the surround- ing scene qualitatively. These results suggest that VLM-assisted segment annotation is a practical route to data-efficient, spatially controllable satellite image editing.

补充信息

↑