发表机构
The University of Texas at Austin; University of Colorado Boulder(德克萨斯大学奥斯汀分校; 科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于CLIP的多尺度补丁驱动模型ReLIQS,解决无参考图像质量评估的分辨率适配、计算效率等问题,在多类基准上泛化能力优于主流基线且成本相当或更低。
AI 中文摘要
无参考图像质量评估(NR IQA)近年受益于深度多模态模型,但多数SOTA系统仍至少违背一项基本要求:要么通过激进调整尺寸丢弃关键质量线索,无法跨分辨率泛化,无法在具有不匹配MOS尺度的异构IQA数据集上联合训练,或计算成本过高。我们提出ReLIQS,即面向图像质量的分辨率无关学习模型(ReLIQS),该模型具备分辨率无关性、保留原始分辨率质量线索、可从多项主观研究学习,且计算高效、预算自适应。ReLIQS是基于CLIP的多尺度补丁驱动架构,同时学习“关注何处”与“如何判断”质量:在包括原始分辨率在内的多种分辨率上采样固定尺寸补丁,并用CLIP视觉主干编码;轻量感知重要性估计器随后预测IQA特定重要性图以选择少量信息补丁,潜在质量轴模块将这些补丁的嵌入聚合为单张图像级分数。在涵盖不同分辨率与失真的真实、合成及AIGC基准上,ReLIQS的泛化能力优于强大的CNN、CLIP及MLLM基线,且计算成本相当或更低。
英文摘要
No-reference image quality assessment (NR IQA) has recently benefited from deep and multimodal models, yet many SOTA systems still violate at least one basic requirement: they either discard critical quality cues via aggressive resizing, fail to generalize across resolutions, cannot be jointly trained on heterogeneous IQA datasets with mismatched MOS scales, or require prohibitive computation. We present \textbf{ReLIQS}, a model for \textbf{Re}solution-agnostic \textbf{L}earning for \textbf{I}mage \textbf{Q}uality with \textbf{S}aliency, which is resolution-agnostic, preserves original-resolution quality cues, learns from multiple subjective studies, and remains computationally efficient and budget-adaptive. ReLIQS is a CLIP-based multiscale patch-driven architecture that learns both \emph{where to look} and \emph{how to judge} quality. Fixed-size patches are sampled across multiple resolutions, including the original resolution, and encoded with a CLIP vision backbone. A lightweight Perceptual Importance Estimator then predicts IQA-specific importance maps to select a small set of informative patches, and a Latent Quality Axis Module aggregates their embeddings into a single image-level score. Across authentic, synthetic, and AIGC benchmarks spanning diverse resolutions and distortions, ReLIQS generalizes better than strong CNN-, CLIP-, and MLLM-based baselines with matching or reduced computational cost.
CommentsAccepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026