arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散目标,保留其标签:通过VLM构建的3D植被场景从少量未标记照片中整理检测器训练数据

Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

Mario Malizia, Marnix Enting, Rob Haelterman, Ken Hasselmann

arXiv 2608.09691首次发表:更新:

发表机构

Royal Military Academy; KU Leuven; Flanders Make(皇家军事学院; 鲁汶大学; 佛兰德制造研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对植被中小物体标记图像稀缺导致检测器跨站点泛化差的问题,通过VLM构建3D植被场景合成标记训练图像,实现无监督站点适应,在排雷基准上性能优于传统跨站点标签复用。

AI 中文摘要

隐藏在植被中的小物体的标记图像稀缺,基于这些图像训练的检测器跨站点泛化能力较差。我们没有重复使用在另一站点收集的标记,而是从部署站点本身的少量未标记照片中合成标记的训练图像。视觉-语言模型(VLM)从一张照片生成粗糙的3D植被场景;将3D物体网格放置在场景中,可直接从场景几何结构得到边界框、分割掩码和每个实例的遮挡情况,无需手动标注。在照片上微调的轻量适配器会对重新渲染的图像进行扩散处理,分级掩码锁设置扩散可接触物体本身的程度。在我们的实验中,该分级是最具影响力的整理选择:轻度扩散物体比完全保护其像素能提升少数类召回率,而无限制的扩散会使物体消失。在人道主义排雷基准测试中,基于这些图像训练的标准检测器在不同随机种子下,始终与在另一站点更大的真实图像标记数据集上训练的对应检测器表现相当或更优;该比较是从少量照片进行无监督站点适应与传统跨站点标签复用的对比。在我们的消融实验中,增益对照片和裁剪预算基本不敏感,且域内准确率无法预测跨站点性能。

英文摘要

Labeled images of small objects hidden in vegetation are scarce, and detectors trained on them generalize poorly across sites. Rather than reusing labels collected at another site, we synthesize labeled training images from a handful of unlabeled photographs of the deployment site itself. A vision--language model generates a coarse 3D vegetation scene from one photograph; placing 3D object meshes in the scene yields bounding boxes, segmentation masks, and per-instance occlusion directly from the scene geometry, without manual annotation. A lightweight adapter fine-tuned on the photographs conditions a diffusion pass that re-textures the renders, and a graded mask-lock sets how much diffusion may touch the object itself. In our runs this grade was the most influential curation choice: lightly diffusing the object improves minority-class recall over fully protecting its pixels, while unrestricted diffusion dissolves it. Trained on these images, a standard detector matched or exceeded its counterpart trained on a larger labeled dataset of real images from a different site, consistently across seeds on a humanitarian-demining benchmark; the comparison is thus unsupervised site adaptation from a handful of photographs against conventional cross-site label reuse. In our ablations the gains were largely insensitive to the photograph and crop budgets, and in-domain accuracy did not predict cross-site performance.

CommentsAccepted at the Curated Data for Efficient Learning (CDEL) Workshop @ ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑