发表机构
TNO - Defence, Security and Safety; Radboud University(TNO国防、安全与安全研究院; 拉德堡德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究以伪装军用车辆检测为场景,利用Qwen Image Edit 2509、Flux.2 Dev等生成式AI模型合成伪装数据,结合LoRA微调提升目标检测器的领域迁移鲁棒性,在植被、网布伪装场景实现显著mAP提升。
AI 中文摘要
目标检测器在领域迁移(如光照、天气或遮挡的变化)下性能往往会下降。这些迁移会改变物体外观,暴露出模型对从训练分布中学到的视觉捷径的依赖,而这些捷径无法跨领域泛化。在数据有限的专业场景中,获取足够的真实世界样本以捕捉此类领域变化尤为困难。基于扩散的生成式图像编辑的最新进展,已通过合成数据增强展现出提升目标检测器域内性能的潜力,但其提升域外鲁棒性的潜力在很大程度上仍未被探索。我们假设生成式图像编辑可在训练数据中模拟可控的领域迁移,有效弥合源域与目标域之间的差距。为验证这一点,我们以伪装军用车辆检测作为具有挑战性的领域迁移场景展开研究。在未伪装数据上训练的检测器,在包含15类车辆的近距地面视角图像(带有植被、网布和多光谱伪装)的真实测试图像上表现出显著性能下降。我们使用两款基于扩散的编辑模型Qwen Image Edit 2509和Flux.2 Dev,在训练数据中合成添加伪装,同时采用经LoRA微调的Qwen版本;以非生成式的黑条遮挡基线作为增强质量的下界。使用在真实与合成数据上训练的GroundingDINO检测器,生成式伪装增强使植被伪装的mAP提升20.1,网布伪装的mAP提升14.4;生成多光谱伪装更具挑战性,但LoRA微调较未伪装基线的性能提升了4.4 mAP。
英文摘要
Object detectors often degrade under domain shifts such as changes in lighting, weather, or occlusion. These shifts alter object appearance and expose a reliance on visual shortcuts learned from the training distribution that do not generalize across domains. Acquiring sufficient real-world samples to capture such domain variation is particularly difficult in specialized, low-data settings. Recent advances in diffusion-based generative image editing have shown promise for improving the in-domain performance of object detectors through synthetic data augmentation. However, their potential to improve out-of-domain robustness remains largely unexplored. We hypothesize that generative image editing can simulate a controlled domain shift in training data, effectively bridging the gap between source and target domains. To test this, we studied camouflaged military vehicle detection as a challenging domain shift scenario. Detectors trained on uncamouflaged data demonstrate substantial degradation on real test imagery containing foliage, netting, and multi-spectral camouflage across 15 vehicle classes in close-up, ground-level imagery. We used two diffusion-based editing models, Qwen Image Edit 2509 and Flux.2 Dev, to synthetically add camouflage to the training data, alongside a LoRA fine-tuned version of Qwen. A non-generative black-bar occlusion baseline served as a lower bound on augmentation quality. Using a GroundingDINO detector trained on real and synthetic data, generative camouflage augmentation yielded substantial mAP improvements for foliage (+20.1) and netting (+14.4) camouflage. Generating multi-spectral camouflage proved more challenging, but LoRA fine-tuning improved performance by 4.4 mAP over the uncamouflaged baseline.
CommentsSubmitted to SPIE Sensors + Imaging 2026