发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory(上海交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LLaVA-Assessor,一个用于视觉质量评估的统一数据构建与模型训练系统,通过质量解释与评分双任务、自适应架构及提示解耦策略,在多个基准上取得优越性能。
AI 中文摘要
与人类视觉系统(HVS)在感知和评估视觉信号质量方面保持一致,是基于机器视觉的视觉质量评估系统的核心目标。随着大型多模态模型(LMMs)的快速发展,视觉问答为在多模态和多任务场景下构建统一的视觉质量评估基础模型提供了一种有前景的范式。受基于HVS的质量评估中经典“感知-决策”过程的启发,我们将基于LMM的机器视觉的视觉质量评估表述为两个互补的任务:“质量解释”和“质量评分”。围绕这些目标,我们提出了LLaVA-Assessor,一个统一的数据构建和模型训练系统。为了支持多模态输入,我们设计了一个自适应模型架构,能够高效处理图像和视频。在数据构建方面,我们制定了严格的人工标注协议和一种新颖的以机器合成主导的数据扩展流程,以构建大规模、高质量的数据集。此外,我们引入了一种简单而有效的提示解耦策略,以缓解多任务学习中的训练目标混淆,从而实现稳定且连贯的联合训练。由此产生的全能型LMM LLaVA-Assessor-GIGA在11个图像/视频质量评分测试集和4个视觉质量解释基准上取得了优越的性能。大量结果表明,将结构化数据构建、自适应模型设计和多任务联合训练相结合,对于自动化视觉质量评估是有效的。我们的工作为开发用于自动视觉质量评估的基础LMM提供了令人信服的见解。项目页面见该https URL。
英文摘要
Aligning with the human visual system~(HVS) in perceiving and evaluating the quality of visual signals is a central objective of machine-vision-based visual quality assessment systems. With the rapid progress of large multi-modal models~(LMMs), visual question answering provides a promising paradigm for building unified foundation models for visual quality assessment under multi-modal and multi-task scenarios. Inspired by the classical ``perception-decision" process in HVS-based quality evaluation, we formulate visual quality assessment for LMM-based machine vision as two complementary tasks: ``quality interpretation'' and ``quality scoring". Centered on these objectives, we propose LLaVA-Assessor, a unified data construction and model training system. To support multi-modal inputs, we design an adaptive model architecture that enables efficient processing of both images and videos. For data construction, we develop rigorous human annotation protocols and a novel machine-synthesis-dominated data expansion pipeline to build a large-scale and high-quality datasets. Furthermore, we introduce a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training. The resulting all-in-one LMM LLaVA-Assessor-GIGA achieves superior performance on $11$ image/video quality scoring test sets and 4 visual quality interpretation benchmarks. Extensive results demonstrate the effectiveness of integrating structured data construction, adaptive model design, and multi-task joint training for automated visual quality assessment. Our work provides compelling insights for developing foundation LMMs for automatic visual quality assessment. Project page at https://github.com/jzhws/LLaVA-Assessor.