发表机构
Université de Mons; Université Polytechnique Hauts-de-France(蒙斯大学; 上法兰西理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统比较Unity仿真与扩散生成两种数据增强范式,发现混合少量合成数据(90%真实+10%合成)能显著提升目标检测性能,最佳mAP@0.5达62.68%,而过量替代会导致域漂移。
AI 中文摘要
现代计算机视觉模型在大型标注数据集上训练时能够实现高精度。在施工安全监控等关键领域,数据收集成本高昂、存在危险且受到伦理约束。本文提出了一项系统性研究,比较两种互补的数据生成范式:(1)基于Unity仿真的渲染和(2)可控扩散生成(CIA),用于真实数据稀缺条件下的目标检测。一个统一的实验框架能够实现真实、仿真和生成数据源的受控数据集混合,同时保持相同的模型和训练设置。使用精确率、召回率、mAP和自定义$\Delta$-指标进行的定量评估表明,单独使用仿真或生成增强均无法实现最优迁移性。仅使用Unity训练相对于真实数据导致mAP@0.5下降$-50\\%$,而仅使用CIA训练显示出较温和的$-16.5\\%$退化。混合组合显著提升性能,其中90\\%真实+10\\%Unity配置达到最佳整体mAP@0.5为$62.68\\%$(较基线提升$+7.64\\%$),而90\\%真实+10\\%CIA配置在精确率上达到最大值$74.45\\%$。结果表明,有限的合成数据纳入增强了泛化能力,而过度的替代则导致域漂移。
英文摘要
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom $Δ$-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of $-50\%$ relative to real data, while CIA-only training shows a milder $-16.5\%$ degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall mAP@0.5 of $62.68\%$ ($+7.64\%$ over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at $74.45\%$. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.