发表机构
Center for Advanced Systems Understanding (CASUS); Helmholtz-Zentrum Dresden-Rossendorf e. V. (HZDR); Institute of Computer Science, University of Wrocław; Cluster of Excellence Physics of Life, TU Dresden(高级系统理解中心; 德累斯顿-罗森多夫亥姆霍兹中心; 弗罗茨瓦夫大学计算机科学研究所; 德累斯顿工业大学生命物理卓越集群)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对科学应用中深度学习计算机视觉数据收集标注难题,提出GraNatPy包用度量引导合成图像数据渲染,经实验证明其能提升渲染数据集质量及检测模型性能,还将数据渲染转化为智能技能优化参数。
AI 中文摘要
科学应用中的深度学习计算机视觉需要通过费力、昂贵且容易出错的过程来收集和注释大型数据集。通过3D建模和渲染生成合成数据可以简化此过程,并通过以编程方式生成注释来提高注释的准确性。然而,在视觉上最小化真实图像和合成图像之间的域差距是主观的,并且缺乏系统的定量指导。我们提出了GraNatPy,一个带有度量的Python包,用于指导渲染场景的改进。我们表明,渲染数据集的真实感、多样性和大小的可量化增加与场景的视觉感知改善和对象检测模型的更高零样本性能相关。此外,我们使用病毒学噬菌斑测定的照片证明,梯度相似性会影响小物体检测的性能,通过混合真实数据和合成数据可以提高性能。最后,我们将程序数据渲染转变为一种智能技能(SynthClaw),以自动进行程序参数优化。
英文摘要
Deep learning computer vision for scientific applications requires collecting and annotating large datasets in a laborious, expensive and error-prone process. Synthetic data generation through 3D modelling and rendering may simplify this process and increase the accuracy of annotations by generating them programmatically. However, minimising the domain gap between real and synthetic images visually is subjective and lacks systematic quantitative guidance. We present GraNatPy, a Python package with metrics to guide improvement of the rendered scene. We show that quantifiable increase in realism, diversity and size of rendered dataset correlates with improved visual perception of the scene and higher zero-shot performance of an object detection model. Furthermore, we demonstrated using photographs of virological plaque assays that gradient similarity affects performance on small object detection, which can be improved by mixing real and synthetic data. Finally, we turn procedural data rendering into an agentic skill (SynthClaw) to automate the procedural parameter optimisation.
Comments17 pages, 3 figures, 4 pages