发表机构
School of Digital and Physical Science, University of Hull; Faculty of Electrical Engineering, University of Sarajevo(数字与物理科学学院,赫尔大学; 电气工程学院,萨拉热窝大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对高性能图像分类模型训练数据不足问题,利用合成数据补充。从多视角分析不同合成图像与真实图像差异,通过实验验证并提出评估应用策略,以提高基于合成图像的分类模型可靠性与安全性。
AI 中文摘要
近年来,高性能图像分类模型在社会上的应用迅速扩展。这些模型需要大量训练数据来提高性能,但获取足够的真实图像往往不切实际,因此合成数据的使用日益广泛。然而,合成图像在训练中未必等同于真实图像。本研究从高维特征空间、颜色空间的低层次统计以及模型训练过程三个角度,系统分析了不同生成方法产生的两类合成图像与真实图像之间的差异。此外,通过考虑实际数据混合场景进行实验验证,提出了对未知质量的合成图像进行初步评估并安全纳入训练的评估和应用策略,旨在提高利用合成图像的图像分类模型的可靠性和安全性。
英文摘要
Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of training data to improve performance, securing sufficient real images is often impractical. As a means to compensate for this shortage, the use of synthetic data is becoming widespread. However, synthetic images are not necessarily equivalent to real images for training purposes. This study systematically analyzes the differences between two types of synthetic images created by different generation methods and real images from three perspectives: high-dimensional feature space, low-level statistics in color space, and the model training process. Furthermore, it experimentally verifies how synthetic data should be utilized by considering realistic data mixing scenarios. This enables the proposal of an evaluation and application strategy for performing preliminary assessments on synthetic images of unknown quality and safely incorporating them into training. This research aims to contribute to enhancing the reliability and safety of image classification models utilizing synthetic images.