arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于低资源土堤检查的砂沸多条件扩散合成

Multi-Conditioned Diffusion Synthesis of Sand Boils for Low-Resource Earthen-Levee Inspection

Padam Jung Thapa, Abdullah Bin Naeem, Ayon Dey, Anav Katwal, Md Tamjidul Hoque

arXiv 2607.08794首次发表:更新:

发表机构

University of Louisiana at Lafayette; Department of Computer Science, Louisiana State University New Orleans(路易斯安那州立大学拉斐特分校; 路易斯安那州立大学新奥尔良分校计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对土堤砂沸像素级检测因注释稀缺受限的问题,提出基于扩散的合成管道,用多分支ControlNet堆栈等技术生成合成图像,经评估各预设利弊后发布标签可靠预设,为低资源土堤检查提供合成图像,限于图像质量等方面,代码等可复现。

AI 中文摘要

土堤上的砂沸是安全关键缺陷,但像素级检测受注释稀缺限制。我们提出一种用于低资源砂沸图像的基于扩散的合成管道。使用通过DreamBooth微调并由多分支ControlNet堆栈条件化的Stable Diffusion XL,该管道从一个小型精选参考集生成合成检查图像。软掩码修复协议保留真实缺陷像素同时重新渲染周围场景。掩码条件化的ControlNet可在选定掩码内生成新的砂沸。文本条件由分类法驱动的提示图集提供。管道生成1020个合成候选图像,815个通过CLIP可接受性过滤器。通过分布和保真度-多样性度量评估图像质量,并审核分布外漂移和记忆情况。没有单一预设占主导,因此发布标签可靠的预设作为默认,并将精选混合物视为自然增强集。我们的工作限于图像质量、标签来源和多样性,下游分割留待未来工作。还发布了代码和工件清单以实现可重复性。

英文摘要

Sand boils on earthen levees are safety-critical defects, but pixel-level detection is limited by scarce annotations. We present a diffusion-based synthesis pipeline for low-resource sand-boil imagery. Using Stable Diffusion XL fine-tuned with DreamBooth and conditioned by a multi-branch ControlNet stack, the pipeline generates synthetic inspection images from a small curated reference set. A soft-mask inpainting protocol preserves the real defect pixels while re-rendering the surrounding scene, avoiding seams and color shifts from prior seamless-cloning compositing. A mask-conditioned ControlNet can also generate a new boil inside a chosen mask, making the mask the segmentation label by construction; however, because large-scale label certification remains unresolved with the available real-trained gate, we release the soft-mask preset as the default. Text conditioning is supplied by a taxonomy-driven Prompt Atlas that expands one domain specification into a stratified, CLIP-validated prompt bank and transfers to new defect classes without code changes. From the real training images, the pipeline produces 1,020 synthetic candidates, of which 815 pass a CLIP admissibility filter. We evaluate image quality using distributional and fidelity-diversity measures against the real reference set and a Poisson baseline, and audit for out-of-distribution drift and memorization. No single preset dominates; each trades off fidelity, diversity, and label reliability. We therefore release the label-reliable preset as the default and treat a curated mixture as the natural augmentation set. Our claims are limited to image quality, label provenance, and diversity; downstream segmentation is left for future work. Code and an artifact manifest are released for reproducibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑