发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出跨模态感知对齐适配器(CMPA),通过感知敏感特征提取器与残差增强降采样策略,解决CLIP在无参考图像质量评估中感知信号被语义淹没的问题,显著超越现有方法。
AI 中文摘要
利用像CLIP这样的大型视觉语言模型最近为无参考图像质量评估(NR-IQA)设立了新的基准。然而,CLIP的对比预训练本质上优先考虑语义不变性,这往往会抑制微妙的感知信号,我们将这一现象称为感知淹没。此外,标准的预处理技术(例如裁剪和插值)进一步加剧了关键高频质量线索的丢失。在本文中,我们提出了跨模态感知对齐适配器(CMPA),这是一个流形感知框架,旨在将感知失真从主导语义中解耦。CMPA引入了一个感知敏感特征提取器(PFE),将CLIP特征投影到一个紧凑的低维子空间中,明确放大由失真引起的离流形偏差。随后,一个跨模态感知对齐注入器(PAI)将这些特征与质量感知的文本锚点对齐,并将它们重新注入到骨干网络中。为了确保输入保真度,我们还设计了一种残差增强感知降采样策略,该策略使用基于恰可察觉差异(JND)引导的频率重注入来自适应地补偿由分辨率引起的信息损失。在多个基准数据集上的广泛评估表明,我们的方法显著优于最先进的方法,有效恢复了淹没在语义密集表示中的感知信号。
英文摘要
Leveraging Large Vision-Language Models like CLIP has recently set new benchmarks for No-Reference Image Quality Assessment (NR-IQA). However, the contrastive pretraining of CLIP inherently prioritizes semantic invariance, which often suppresses subtle perceptual signals, a phenomenon we term perceptual submergence. Furthermore, standard preprocessing techniques (e.g., cropping and interpolation) further exacerbate the loss of critical high-frequency quality cues. In this paper, we propose the Cross-modal Perception Alignment Adapter (CMPA), a manifold-aware framework designed to disentangle perceptual distortions from dominant semantics. CMPA introduces a Perception-Sensitive Feature Extractor (PFE) that projects CLIP features into a compact, low-dimensional subspace, explicitly magnifying distortion-induced off-manifold deviations. Subsequently, a Cross-Modal Perception Alignment Injector (PAI) aligns these features with quality-aware text anchors and re-injects them into the backbone. To ensure input fidelity, we also devise a Residual-enhanced Perceptual Downscaling strategy that adaptively compensates for resolution-induced information loss using Just Noticeable Difference (JND) guided frequency re-injection. Extensive evaluations on several benchmark datasets demonstrate that our approach significantly outperforms state-of-the-art methods, effectively recovering the perceptual signals submerged in semantic-dense representations.
CommentsAccepted by ICML2026