arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有补丁都一样:可见-红外预训练中的采样很重要

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

Qiwei Ma, Bin Deng, Junjie Zhu, Qiangjuan Huang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li

arXiv 2607.20238首次发表:更新:

发表机构

Yuelushan Center for Industrial Innovation; Intelligent Game and Decision Lab(岳麓山工业创新中心; 智能游戏与决策实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对可见-红外预训练,提出重要性感知采样(IAS)方法,通过从红外结构线索导出权重、学习重要性掩码及采用补丁课程学习策略调整训练重点,在多个VIS-IR基准实验中比基线有改进。

AI 中文摘要

可见-红外(VIS-IR)对齐是强大的多传感器感知的关键预训练任务。大多数现有方法使用均匀的逐补丁对比学习,但在VIS-IR数据中这可能不可靠,因为成像物理差异使一些空间配对区域本质上可比性较低,同等强度对齐会阻碍表示学习和下游迁移。本文从采样角度重新审视VIS-IR预训练,提出重要性感知采样(IAS),它基于补丁可靠性调整训练重点。具体包括从红外结构线索导出补丁权重来重新加权对比目标;用轻量级采样器学习软重要性掩码,可从手工先验热启动;采用从高可靠性区域到更难补丁逐步扩展的补丁课程学习策略。IAS是即插即用的,与逐补丁/相关性级对齐和图像级对比基线都兼容。在多个VIS-IR基准上的广泛实验表明,相对于强大基线有持续改进,包括红外语义分割、红外目标检测、VIS语义分割和跨模态检索任务。代码将在该https网址发布。

英文摘要

Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can be unreliable in VIS-IR data because imaging-physics differences make some spatially paired regions inherently less comparable, and aligning them with equal strength hinders representation learning and downstream transfer. In this paper, we revisit VIS-IR pre-training from a sampling perspective and propose Importance-Aware Sampling (IAS), which adjusts training emphasis based on patch reliability. Specifically, IAS (i) derives patch weights from infrared structural cues and uses them to reweight the contrastive objective; (ii) learns a soft importance mask with a lightweight sampler, optionally warm-started from the hand-crafted prior; and (iii) employs a patch curriculum learning strategy that gradually expands from high-reliability regions to harder patches. It is worth noting that IAS is plug-and-play and works with both patch-/correlation-level alignment (e.g., UNIV-style) and image-level contrastive baselines (e.g., ImageBind-style). Extensive experiments on multiple VIS-IR benchmarks demonstrate consistent improvements over strong baselines, including for IR semantic segmentation, IR object detection and VIS semantic segmentation and cross-modal retrieval task. Code will be released on https://github.com/KlayMa527/IAS.

Comments13 pages, 11 figures,

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑