发表机构
Yuelushan Center for Industrial Innovation; Intelligent Game and Decision Lab(岳麓山工业创新中心; 智能游戏与决策实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对可见-红外预训练,提出重要性感知采样(IAS)方法,通过从红外结构线索导出权重、学习重要性掩码及采用补丁课程学习策略调整训练重点,在多个VIS-IR基准实验中比基线有改进。
AI 中文摘要
可见-红外(VIS-IR)对齐是强大的多传感器感知的关键预训练任务。大多数现有方法使用均匀的逐补丁对比学习,但在VIS-IR数据中这可能不可靠,因为成像物理差异使一些空间配对区域本质上可比性较低,同等强度对齐会阻碍表示学习和下游迁移。本文从采样角度重新审视VIS-IR预训练,提出重要性感知采样(IAS),它基于补丁可靠性调整训练重点。具体包括从红外结构线索导出补丁权重来重新加权对比目标;用轻量级采样器学习软重要性掩码,可从手工先验热启动;采用从高可靠性区域到更难补丁逐步扩展的补丁课程学习策略。IAS是即插即用的,与逐补丁/相关性级对齐和图像级对比基线都兼容。在多个VIS-IR基准上的广泛实验表明,相对于强大基线有持续改进,包括红外语义分割、红外目标检测、VIS语义分割和跨模态检索任务。代码将在该https网址发布。
英文摘要
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can be unreliable in VIS-IR data because imaging-physics differences make some spatially paired regions inherently less comparable, and aligning them with equal strength hinders representation learning and downstream transfer. In this paper, we revisit VIS-IR pre-training from a sampling perspective and propose Importance-Aware Sampling (IAS), which adjusts training emphasis based on patch reliability. Specifically, IAS (i) derives patch weights from infrared structural cues and uses them to reweight the contrastive objective; (ii) learns a soft importance mask with a lightweight sampler, optionally warm-started from the hand-crafted prior; and (iii) employs a patch curriculum learning strategy that gradually expands from high-reliability regions to harder patches. It is worth noting that IAS is plug-and-play and works with both patch-/correlation-level alignment (e.g., UNIV-style) and image-level contrastive baselines (e.g., ImageBind-style). Extensive experiments on multiple VIS-IR benchmarks demonstrate consistent improvements over strong baselines, including for IR semantic segmentation, IR object detection and VIS semantic segmentation and cross-modal retrieval task. Code will be released on https://github.com/KlayMa527/IAS.
Comments13 pages, 11 figures,