发表机构
TU Berlin; University of Stuttgart; TU Braunschweig(柏林工业大学; 斯图加特大学; 布伦瑞克工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低图像预算下的工业目标检测,提出语义引导的域随机化(S-GDR),结合VLM语义描述与扩散背景合成,在200张图像预算下达到mAP50-95=0.739,优于现有基线。
AI 中文摘要
在高混合、低产量(HMLV)汽车制造中,重新训练视觉感知流程必须在严格的标注、能源和时间预算下进行,然而大多数合成数据生成(SDG)策略仍需要数千张图像。本研究评估了语义引导的域随机化(S-GDR),这是一种无需标注的自适应流程,它将基于视觉语言模型(VLM)的语义描述(针对小型未标注真实参考集)与基于扩散的背景合成(由ControlNet和IP-Adapter条件化的Stable Diffusion XL(SDXL))以及基于掩码的对象合成相结合。在汽车多目标检测基准上,使用固定的200张合成训练图像预算,S-GDR在真实保留测试集上达到了mAP50-95 = 0.739,优于域随机化渲染基线(mAP50-95 = 0.697),以及亮度过滤、感知哈希、CycleGAN风格迁移和共享相同200张图像预算的无引导扩散变体。这些初步观察将S-GDR定位为极端数据稀缺场景下一种有前景的免标注替代方案。
英文摘要
Retraining visual perception pipelines in High-Mix, Low-Volume (HMLV) automotive manufacturing must be carried out under tight annotation, energy, and time budgets, yet most Synthetic Data Generation (SDG) strategies still operate in the thousands of images. This work evaluates Semantically-Guided Domain Randomization (S-GDR), an annotation-free adaptation pipeline that couples Vision-Language Model (VLM)-based semantic captioning of a small unannotated real reference set with diffusion-based background synthesis (Stable Diffusion XL (SDXL) conditioned by ControlNet and IP-Adapter) and mask-based object composition. On an automotive multi-object detection benchmark and with a fixed budget of 200 synthetic training images, S-GDR reaches mAP50-95 = 0.739 on a real held-out test set, outperforming a domain-randomized render baseline (mAP50-95 = 0.697) as well as brightness filtering, perceptual hashing, CycleGAN style transfer, and unguided diffusion variants sharing the same 200-image budget. These initial observations position S-GDR as a promising annotation- free alternative for extreme data-scarcity regimes.