arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SC-Diff:用于可见光到红外图像转换的语义校准扩散模型

SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation

Junyin Zhang, Siyu Huang, Jianxiong Ye, Haowei Gong, Ruicheng Zhang, Deyu Meng, Chenqiang Gao

arXiv 2608.08555首次发表:更新:

发表机构

School of Intelligent Systems Engineering, Shenzhen Campus of Sun Yat-sen University; Tsinghua Shenzhen International Graduate School, Tsinghua University; School of Mathematics and Statistics, Xi’an Jiaotong University(中山大学深圳校区智能系统工程学院; 清华大学深圳国际研究生院; 西安交通大学数学与统计学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SC-Diff 是一种语义校准潜扩散框架,通过结合语义先验作为条件与自注意力校准,提升可见光转红外图像的语义一致性,生成更适用于红外目标检测的合成训练数据。

AI 中文摘要

可见光到红外图像转换为利用丰富的可见光图像扩展红外训练数据提供了实用途径,扩散模型因强大的生成性能在该任务中颇具潜力。然而,现有基于扩散的方法通常仅将语义先验作为外部条件,未显式约束去噪网络内部的 token 交互,因此难以保留可靠标注复用所需的目标位置、形状和语义布局。我们提出 SC-Diff,一种语义校准的潜扩散框架,该框架将语义先验同时用于条件引导和内部自注意力校准。首先,带有预定义文本提示的预训练 SAM3 模型从可见光图像中提取类别特定的语义掩码;这些掩码被合并成语义图并与可见光图像融合作为输入条件。同一张图被转换为 token 级语义标签以校准去噪网络中的自注意力。基于这些标签,我们引入语义引导自注意力校准(SGSC),其为同一类别的查询-键对自适应施加正偏置。逐查询的校准强度取决于注意力在语义类别间的分散程度以及分配给该查询自身类别的注意力,原始注意力分数进一步调制该偏置,对响应更强的同类别键给予更大校准。这种软校准减少了跨类别干扰,同时保留了全局上下文交互,从而提升生成红外图像的语义一致性。大量实验表明,SC-Diff 提升了感知质量,并为下游红外目标检测生成了更有效的合成训练数据。

英文摘要

Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are promising for this task because of their strong generative performance. However, existing diffusion-based methods typically use semantic priors only as external conditions, without explicitly regulating token interactions within the denoising network. Consequently, they struggle to preserve object locations, shapes, and semantic layouts required for reliable annotation reuse. We propose SC-Diff, a semantically calibrated latent diffusion framework that uses semantic priors for both conditional guidance and internal self-attention calibration. A pretrained SAM3 model with predefined text prompts first extracts category-specific semantic masks from visible images. These masks are merged into a semantic map and fused with the visible image as the input condition. The same map is converted into token-level semantic labels to calibrate self-attention in the denoising network. Based on these labels, we introduce Semantic-Guided Self-Attention Calibration (SGSC), which adaptively applies positive biases to query-key pairs of the same category. The query-wise calibration strength depends on the dispersion of attention across semantic categories and the attention assigned to the query's own category. The original attention scores further modulate the bias, giving greater calibration to same-category keys with stronger responses. This soft calibration reduces cross-category interference while retaining global contextual interactions, thereby improving semantic consistency in generated infrared images. Extensive experiments show that SC-Diff improves perceptual quality and produces more effective synthetic training data for downstream infrared object detection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑