发表机构
Dalian University of Technology; New York University; Nanjing University of Science and Technology; University of California, San Francisco(大连理工大学; 纽约大学; 南京理工大学; 加利福尼亚大学旧金山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态重识别中文本噪声与模态结构差异问题,提出基于正激励噪声和结构化提示调制的框架,通过语义跨模态调制、结构感知提示适配和上下文感知稀疏融合,在三个基准上验证了有效性与鲁棒性。
AI 中文摘要
多模态目标重识别(ReID)受益于异构成像模态之间的互补信息。为了进一步丰富语义表示,最近文本描述被作为额外模态引入。然而,现有的视觉-语言方法通常将文本描述视为干净、确定性的信号,忽视了其固有的噪声,包括模态不匹配的短语和语义模糊的表达。此外,现有方法缺乏显式机制来协调模态之间细粒度的结构差异,即使在高层次语义对齐之后也是如此。为解决这些挑战,我们提出了一种以正激励噪声(π-噪声)和结构化提示调制为核心的新框架。首先,语义跨模态调制器利用任务感知的π-噪声(从以视觉和文本输入为条件的分布中采样)来扰动全局标记,实现语义引导的跨模态补偿。其次,结构感知提示适配器通过提示注入可学习的几何先验以增强空间一致性。第三,上下文感知稀疏融合模块提炼结构上下文以指导自适应融合,同时保护身份特征免受噪声局部细节的影响。在三个多模态ReID基准上的实验证明了我们方法的有效性和鲁棒性。代码可在该https URL获取。
英文摘要
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as clean, deterministic signals and overlook their inherent noise, including modality-mismatched phrases and semantically ambiguous expressions. Moreover, prevailing methods lack explicit mechanisms to reconcile fine-grained structural discrepancies between modalities, even after high-level semantic alignment. To address these challenges, we propose a novel framework centered on Positive-Incentive Noise (π-noise) and structured prompt modulation. First, the Semantic Cross-Modal Modulator harnesses task-aware π-noise, sampled from a distribution conditioned on both visual and text inputs, to perturb global tokens and enable semantics-guided cross-modal compensation. Second, the Structure-Aware Prompt Adapter injects learnable geometric priors via prompts to enhance spatial consistency. Third, the Context-Aware Sparse Fusion module distills structural context to guide adaptive fusion while shielding identity features from noisy local details. Experiments on three multi-modal ReID benchmarks demonstrate the effectiveness and robustness of our approach. The code is available at https://github.com/zw-absin/INSPI.
CommentsAccepted by ECCV 2026. The version of record may differ slightly