面向交通标志增强的、具备物理一致性的结构化先验引导扩散修复
Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation
浏览论文内容
中文总结 AI 辅助
针对交通标志检测的长尾数据问题,提出结构化先验引导的物理一致性扩散修复框架,在多指标上优于7种对比方法,合成数据显著提升稀有标志检测性能
中文摘要 AI 辅助
交通标志检测面临长尾数据分布问题。从监管角度来看,许多稀有标志与常见标志同等重要,但它们的样本数量极少。生成式数据增强是解决该问题的一种方法。然而,将通用修复模型直接应用于标志区域时,会出现数字扭曲、几何与透视变形以及颜色偏移等问题。我们将此归因于一个单一的缺陷:条件信号对于标志的物理构成而言过于抽象。我们提出了一种具备物理一致性的结构化先验引导扩散修复框架,该框架通过三条正交路径注入标志的语义、外观和几何先验:JSON格式的文本提示、经测量的主色渲染的正面矢量模板(通过IP-Adapter实现)、仿射对齐的矢量模板(通过ControlNet实现)。两项物理一致性损失分别通过CIELAB色度L1项约束颜色,通过Sobel梯度项约束边缘结构。我们在AMAP内部收集的大量图像上进行自监督重构训练,随后在不同来源的公开TT100K-2021数据集上进行零样本评估。该方法采用约14亿参数的Stable Diffusion 1.5主干模型,在重构保真度、物理一致性和语义可控性的所有指标上均优于7种代表性对比方法。其OCR精确匹配率达到91.1%,而120亿参数的工业模型FLUX.1 Fill [dev]仅为44.2%,且该方法仅需FLUX.1 Fill [dev]模型1/14的推理时间。留一法消融实验证实,三条先验路径和两项损失项各自均有贡献。在下游检测任务中,合成数据使稀有类别的组池化AP50较仅使用真实数据的基准提升了1.23倍至7.40倍。代码和预训练模型可在该https URL获取。
英文摘要
Traffic sign detection faces a long-tailed data distribution. Many rare signs matter as much as common ones from a regulatory standpoint, yet they have very few samples. Generative data augmentation is one way out. General-purpose inpainting models, however, distort digits, deform geometry and perspective, and shift colours when applied directly to sign regions. We trace this to a single gap: the conditioning signal is too abstract for the physical composition of a sign. We propose a structured-prior-guided diffusion inpainting framework with physical consistency. It injects the semantic, appearance and geometric priors of a sign through three orthogonal pathways: a JSON-formatted text prompt, a front-view vector template rendered with measured dominant colours (via IP-Adapter), and an affine-aligned vector template (via ControlNet). Two physical consistency losses constrain colour with a CIELAB chromaticity $L_1$ term and edge structure with a Sobel gradient term. We train by self-supervised reconstruction on a large set of images collected in-house at AMAP, then evaluate zero-shot on the public TT100K-2021 dataset, a different source. Our method uses a Stable Diffusion 1.5 backbone of about 1.4B parameters. It beats seven representative competitors on every metric of reconstruction fidelity, physical consistency and semantic controllability. Its OCR exact-match rate reaches 91.1\%, against 44.2\% for the 12B industrial model FLUX.1 Fill [dev], and it needs only $1/14$ of that model's inference time. Leave-one-out ablations confirm that each of the three prior pathways and both loss terms contribute on their own. In downstream detection, the synthetic data raises the group-pooled AP50 of rare classes by $1.23\times$ to $7.40\times$ over a real-data-only baseline. Code and pre-trained models are available at https://github.com/52hz-whale/TrafficSignInpaint.
发表机构
- AMAP, Alibaba Group(高德地图(阿里巴巴集团旗下))
机构由 AI 辅助整理,请以论文原文为准。