arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24110cs.CVcs.LGphysics.optics

超越融合:用于无校准红外超分辨率和红外-可见光融合的自对齐潜在扩散

BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

Minchong Chen, Xiaoyun Yuan, Minyu Cao, Jianing Zhang, Jun Zhang, Shuyang Liu, Xiaokang Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对移动红外-可见光成像中跨传感器未对准问题,提出超越融合框架,通过引入跨模态自对齐模块及未对准增强模块,实现无校准条件下的红外超分辨率和红外-可见光融合,经实验验证了其有效性。

中文摘要 AI 辅助

移动红外-可见光成像通常将紧凑型红外传感器与高分辨率可见光相机配对以实现互补感知。然而,由不同光学、视角、视野和曝光时间导致的跨传感器未对准阻碍了实际部署。本文提出超越融合,这是一个用于无校准可见光引导红外超分辨率和红外-可见光融合任务的统一潜在扩散框架。该框架支持特定任务训练和联合训练。超越融合在去噪U-Net中引入跨模态自对齐模块,将红外和可见光潜在令牌重新组织到共享注意力空间以学习内容自适应跨模态对应。结合未对准增强模块,该模型能利用可见光结构和语义线索同时保持热一致性,在未校准条件下实现高频红外重建和信息丰富的融合图像生成。在公共基准和移动红外-可见光成像系统上的大量实验表明,该模型在各种输入条件下都有强大性能。消融研究、统一训练分析和下游行人检测进一步验证了超越融合在无校准多模态成像中的有效性。

英文摘要

Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-specific training and joint training where two tasks are optimized and executed as two readouts of the same generative process. Instead of relying on explicit registration or geometric warping, BeyondFusion introduces a cross-modal self-aligning (CMSA) module into the denoising U-Net. CMSA reorganizes infrared and visible latent tokens into a shared attention space to learn content-adaptive cross-modal correspondence during the denoising process. Together with misalignment augmentation module, the model is facilitated to exploit visible structural and semantic cues while preserving thermal consistency, enabling high-frequency infrared reconstruction and informative fused-image generation under uncalibrated conditions. Extensive experiments on public benchmarks and a mobile infrared-visible imaging system show strong performance across aligned inputs, low-resolution infrared observations, synthetic misalignments, and real mobile captures with unsynchronized sensors. Ablation studies, unified training analysis, and downstream pedestrian detection further validate the effectiveness of BeyondFusion for calibration-free multimodal imaging.

补充信息

↑