发表机构
Monash University; Shanghai Jiao Tong University(莫纳什大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出端到端自动化流水线,结合贝塞尔曲线骨架、GAN和双ControlNet扩散框架,实现可控裂缝数据合成,在CRACK500和CrackTree200上优于现有增强方法。
AI 中文摘要
自动化裂缝检测日益依赖深度学习,但其可靠性受到稀缺且可控性差的缺陷数据的限制。现有的生成式增强方法通常将裂缝合成视为通用图像生成任务,对形态、边界保真度和场景上下文提供的控制不足。本文提出了一种用于可控裂缝数据合成的端到端自动化流水线,将裂缝几何和检测上下文形式化为可复用的计算约束。首先,通过程序化采样的贝塞尔曲线骨架利用GAN转换为逼真的裂缝掩码,实现无需手动掩码设计的多样化裂缝形态的可扩展生成。其次,双ControlNet扩散框架将外观引导与几何引导解耦,基于边缘的分支强制执行严格的边界一致性。该框架支持无背景合成和上下文感知修复。在CRACK500和CrackTree200上的实验表明,与现有增强基线相比持续获得提升,展示了用于自动化裂缝检测数据生成的可扩展工程信息学工作流。
英文摘要
Vision-based crack inspection depends on segmentation networks whose reliability depends on the quantity, diversity and label quality of their training data. Pixel-level annotations are costly, and crack images of specific structures are scarce. Generative augmentation can supply additional data, but existing methods address isolated steps. They reuse annotated masks, offer limited control over crack geometry, and adopt the conditioning mask as the label without checking it. This paper presents an end-to-end pipeline that produces labelled crack data without manual annotation and assesses the reliability of these data and of the detectors trained on them. Procedurally sampled Bézier skeletons with guaranteed geometric properties are converted into crack masks by a generative adversarial network (GAN). A dual-ControlNet Stable Diffusion model renders the masks as crack images, either on text-described surfaces or on user-provided backgrounds. An ensemble of segmentation networks trained on real images combines its agreement with the inherited label and its internal disagreement into a pixel-wise label confidence. This confidence weights the training loss instead of removing samples with a threshold. The trained detectors are evaluated with image-space probability of detection (POD) and calibration analyses. On CRACK500 and CrackTree200, the pipeline improves five segmentation networks over conventional, diffusion-based and flow-matching-based augmentation, and on CRACK500 confidence weighting yields a higher accuracy than threshold filtering at every tested threshold. On CRACK500, the crack width that U-Net detects with 90\% probability at 95\% confidence decreases from 8.0 to 4.3 pixels, and the expected calibration error decreases from 14.2\% to 9.6\%.