图像增强作为深度学习图像检索系统的测试生成方法
Image Augmentation as Test Generation for Deep Learning-Based Image Retrieval Systems
AI总结:
该研究将图像增强技术作为深度学习图像检索系统的测试生成方法,通过实证研究筛选出天气模拟和SaSPA等适配的增强技术,为构建有效测试套件提供了实用指南。
AI中文摘要:
确保基于深度学习的图像检索系统的可靠性是软件工程领域的一项挑战。本文有两方面贡献:(1)对增强技术与生成技术进行文献综述,识别出50种技术并将其整理为包含10个类别的分类体系;(2)开展大规模实证研究,评估这些技术作为基于嵌入的图像检索系统测试生成器的效果。增强图像通过Amazon Titan和OpenCLIP进行嵌入处理,从四个分析维度评估:(1)嵌入空间相似度;(2)通过4种估计器测量的嵌入不确定性;(3)由LLaVA评分的语义真实性;(4)检索失败率。实验在三个数据集上进行:CIFAR-10、ImageNet-1K以及工业合作伙伴March Networks提供的数据集。在所有评估的数据集、嵌入模型及每种技术测试的单一严重程度级别下,天气模拟技术和SaSPA是能产生最高嵌入不确定性与失败率,同时在性能稳定性、视觉真实性及增强有效性间保持良好平衡的图像增强/生成技术。本文讨论的结果具有配置特定性,在更温和或更强的扰动设置下可能发生变化。相比之下,基于GAN的增强技术的真实性处于较低水平,表明存在合成伪影与感知不一致问题,降低了其生成真实测试输入的适用性。总体而言,本文的发现为选择能最大化测试多样性同时保留真实图像特征的增强技术提供了实用指南,从而能够构建全面且有效的图像检索系统测试套件,并通过使用变形测试降低手动数据标注的成本。
英文摘要:
Ensuring the reliability of deep learning-based image retrieval systems is a software engineering challenge. This paper presents a dual contribution: (1) a literature review of augmentation and generation techniques which resulted in the identification of 50 techniques which we organized into a ten-category taxonomy, and (2) a large-scale empirical study that evaluates these techniques as test generators for embedding-based image retrieval systems. Augmented images are embedded using Amazon Titan and OpenCLIP, and evaluated across four analytical dimensions: (1) embedding-space similarity, (2) embedding uncertainty measured via four estimators, (3) semantic realism scored by LLaVA, and (4) retrieval failure rate. Experiments are performed on three datasets: CIFAR-10, ImageNet-1K, and a dataset from an industrial partner (March Networks). Across all evaluated datasets and embedding models, and under the single severity level tested for each technique, weather simulation and SaSPA are the image augmentation/generation techniques that produce the highest embedding uncertainty and failure rates while maintaining a favorable balance between performance stability, visual realism, and augmentation effectiveness. The results we discuss are configuration-specific and may shift under milder or stronger perturbation settings. In contrast, GAN-based augmentation techniques are among the lowest in realism, indicating the presence of synthetic artifacts and perceptual inconsistencies that reduce their suitability to produce realistic test inputs. Overall, our findings provide practical guidelines for selecting augmentation techniques that maximize test diversity while preserving realistic image characteristics, thereby enabling the construction of comprehensive and effective test suites for image retrieval systems while reducing the cost of manual data labeling through the use of metamorphic testing.