arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19886cs.CV

MTVDiff:用于增强热红外到可见光面部翻译的多模态条件潜扩散

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

  • School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
  • Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(新一代人工智能技术及其交叉应用重点实验室(东南大学))
  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhiyuan Xia, Haojie Li, Jingyu Lin, Yiguo Qiao, Cunjian Chen

AI总结:

针对热红外到可见光面部翻译的挑战,提出MTVDiff框架,通过双分支交叉注意力融合等三个核心技术,整合深度和文本信息,在多指标上优于现有方法,提升图像质量和面部验证性能,推动跨光谱面部图像翻译发展。

AI中文摘要:

热红外到可见光的面部翻译面临几何不连续、语义属性不匹配和身份退化等基本挑战。我们提出了MTVDiff,一种新颖的多模态潜扩散框架,它协同整合深度和文本信息来解决这些限制,同时保留身份特征。该框架有三个核心技术贡献:用于多尺度热红外-深度特征提取和融合的双分支交叉注意力融合模块;用于语义引导生成的门控文本到视觉特征对齐机制;用于自适应多模态先验整合的空间特征变换。在MCXFace和SpeakingFaces数据集上的大量实验表明,我们的多模态方法在多个指标上显著优于现有的基于GAN和扩散的方法,在图像质量和面部验证性能上都有大幅提升。我们的工作为在不同光照条件下运行的面部识别系统提供了强大的解决方案,并通过有效的多模态整合推动了跨光谱面部图像翻译的技术发展。

英文摘要:

Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.

补充信息

↑