arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09122cs.CVcs.AI

基于大型多模态模型的UGC图像视觉失真检测

Visual Distortion Detection in UGC Images Using Large Multimodal Models

Ziheng Jia, Yingji Liang, Jiaying Qian, Xiongkuo Min

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有LMM图像失真检测方法准确率不足及合成失真泛化差距问题,提出VIGIL模型,构建VIGIL-140K数据集,利用LLM解码器多层级特征检测并缓解FG-BG分离问题,在两类任务上性能优于基线。

中文摘要 AI 辅助

感知质量的局部描述长期以来一直是图像质量评估(IQA)领域中一个至关重要但尚未得到充分探索的挑战。现有的基于大型多模态模型(LMM)的方法主要依赖于文本驱动的监督微调(SFT),然而这种训练范式在检测准确率方面存在明显局限性。此外,通常作为主要训练数据源的合成失真图像在实际场景部署时会出现显著的泛化差距,因此合成到真实(S2A)问题成为一项关键挑战。针对这些问题,我们提出了VIGIL,该模型利用LMM架构实现精准的视觉失真检测。我们从超过100万份样本的候选池中构建了VIGIL-140K训练集,该数据集包含超过14万张失真图像,这些图像经过严格的质量筛选和精心设计的失真注入,涵盖8个主要的合成失真类别。我们的模型利用大型语言模型(LLM)解码器的不同层,将其视为多个检测器,使用多级特征同步执行失真检测;此外,我们保留了分配给非失真类别的预测中的失真线索,这有助于缓解S2A问题中常见的模糊前景-背景(FG-BG)分离问题。经过后处理,我们的模型在域内合成失真检测和S2A任务上均始终优于强基线方法。

英文摘要

The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal models (LMMs) predominantly rely on text-driven supervised fine-tuning (SFT). However, this training paradigm exhibits notable limitations in detection accuracy. Moreover, synthetically distorted images, which are often used as the primary training data source, show a significant generalization gap when deployed in real-world scenarios; thus, the \textbf{synthetic-to-authentic (\textit{S2A})} problem represents a critical challenge. Motivated by these issues, we propose \textbf{\textit{VIGIL}}, which leverages the LMM architecture for precise visual distortion detection. From a candidate pool of over 1000K samples, we construct the \textbf{\textit{VIGIL-140K}} training set, which consists of over 140K distorted images. These images are obtained through rigorous quality filtering and carefully crafted distortion injection, covering 8 major synthetic distortion categories. Our model leverages different layers of the large language model (LLM) decoder, treating them as \textit{multiple detectors} that perform synchronous distortion detection using multi-level features. Additionally, we retain distortion cues from predictions assigned to the non-distortion class, which helps mitigate the ambiguous foreground-background (\textit{FG-BG}) separation commonly encountered in the \textit{S2A} problem. After post-processing, our model consistently outperforms strong baselines on both in-domain synthetic distortion detection and \textit{S2A} tasks.

↑