AI 中文总结
本研究提出QR-STT框架,作为隐蔽黑盒攻击手段,可定向操控IR-VLMs的语义对齐,其扰动还能跨任务迁移,凸显了对该类攻击进行鲁棒性评估的必要性。
AI 中文摘要
红外视觉语言模型(IR-VLMs)将热感知扩展到开放词汇分类、图像字幕和视觉问答任务,但目前针对结构化热扰动的鲁棒性及跨模态语义对齐的稳定性研究尚不充分。我们提出QR结构化热触发装置(QR-STT),这是一种隐蔽、无需训练的黑盒框架,用于对IR-VLMs进行定向语义引导。QR-STT保留QR码的功能区域,同时优化其内部模块,每个模块被分配冷、中性或热的热状态。该框架联合搜索模块拓扑结构和渲染参数,包括位置、尺度、旋转、强度、模糊度和圆度。采用带贪婪模块翻转优化的三阶段无梯度程序,有效处理混合离散与连续的搜索空间。目标函数促进与攻击者选定目标的对齐、抑制源类证据,并对QR结构和视觉相似性进行正则化。在多个CLIP风格编码器上的实验表明,QR-STT可一致地将图像文本对齐重定向到选定概念,同时保持视觉隐蔽性。针对分类优化的扰动也可迁移到图像字幕和VQA任务,在生成输出中导致与目标一致的语义漂移。这些结果表明QR结构化热模式是语言驱动红外感知的可解释攻击面,凸显了针对结构化跨任务语义攻击进行鲁棒性评估的必要性。
英文摘要
Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot thermal state. The framework jointly searches module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. A three-stage gradient-free procedure with greedy module-flip refinement efficiently handles the mixed discrete and continuous search space. The objective promotes alignment with an attacker-selected target, suppresses source-class evidence, and regularizes QR structure and visual similarity. Experiments on multiple CLIP-style encoders show that QR-STT consistently redirects image-text alignment toward chosen concepts while maintaining visual stealth. Perturbations optimized for classification also transfer to image captioning and VQA, causing target-consistent semantic drift in generated outputs. These results identify QR-structured thermal patterns as an interpretable attack surface for language-driven infrared perception and highlight the need for robustness evaluation against structured cross-task semantic attacks.