AI 中文总结
针对复杂退化下红外-可见光图像融合的挑战,提出文本引导的TGFusion框架,通过多流联合流Transformer实现动态文本引导,在多类退化场景下取得优异融合性能。
AI 中文摘要
在真实退化场景下的红外-可见光图像融合是一项具有挑战性的任务,因为退化不仅会导致观测图像中可靠的模态特定信息丢失,还会阻碍融合过程。近期研究表明,文本可提供退化特征的先验信息,补充受损输入图像中有限的可用证据,从而促进融合。然而,现有方法通常将固定的全局文本表示注入视觉特征,使得文本引导难以适应空间变化的退化、局部结构和热显著性。为此,我们提出了TGFusion,一种文本引导的潜在空间流匹配框架,它将退化抑制与跨模态融合相统一。TGFusion将任务、退化和生成线索编码为结构化提示。为充分利用这些先验,我们设计了提示条件多流联合流Transformer,该Transformer将文本表示为与融合、可见光和红外流并列的独立语义流。联合注意力实现了语义和视觉表示之间的令牌级双向交互与逐层更新,使退化语义能够动态引导可靠信息选择和融合潜在生成。在公共基准及复杂退化场景下的大量实验表明,TGFusion在感知质量、图像自然度、结构细节保留和红外显著性保持方面取得了更优或具有竞争力的性能,同时在多种单一及复合退化下保持了鲁棒性。
英文摘要
Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate that text can provide prior information about degradation characteristics, complementing the limited evidence available from corrupted input images and facilitating fusion. However, existing methods typically inject fixed global text representations into visual features, making it difficult for textual guidance to adapt to spatially varying degradations, local structures, and thermal saliency. To this end, we propose TGFusion, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion. TGFusion encodes task, degradation, and generation cues into structured prompts. To fully exploit these priors, we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream alongside fusion, visible, and infrared streams. Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations, allowing degradation semantics to dynamically guide reliable information selection and fusion latent generation. Extensive experiments on public benchmarks and complex degradation scenarios demonstrate that TGFusion achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Comments12 pages, 9 figures, including supplementary material