arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16742cs.CV

人工智能生成的以人为中心的视频的多维质量评估:数据集与模型

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

  • Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所)
  • USC-SJTU Institute of Cultural and Creative Industry, Shanghai Jiao Tong University(上海交通大学南加州大学文化创意产业学院)
  • Polytech Nantes, Université de Nantes(法国南特大学高等理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Sijing Wu, Yunhao Li, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai

AI总结:

研究针对AI生成的以人为中心的视频质量评估问题,提出扩展数据集HVEval+并构建MoE-Rater模型,该模型采用专家混合及三阶段训练策略,能统一多种任务,在相关数据集上性能优越,推动了视频质量评估及T2V模型优化。

AI中文摘要:

人工智能生成的以人为中心的视频在众多现代应用中起着关键作用。然而,它们常存在质量问题和语义不匹配,凸显了对此类视频进行有效质量评估的重要性。为此,我们扩展之前的数据集HVEval,加入成对偏好注释,得到HVEval+,这是用于人工智能生成的以人为中心视频的最大整体质量评估数据集。它包含基于综合分类法的1000个提示、由24个文本到视频(T2V)模型生成的20000个视频以及大量人工注释。同时,我们提出MoE-Rater,一种受专家混合(MoE)启发且基于多模态大语言模型(MLLM)的一体化方法,支持在单个模型内进行多维质量评分、多维成对比较和特定类别问答。具体介绍了投影专家混合(MoPE)和LoRA专家混合(MoLE),以及由任务感知预训练、特定任务适应和自适应路由优化组成的三阶段训练策略。大量实验和综合分析表明HVEval+数据集和MoE-Rater方法在推进人工智能生成视频质量评估及促进T2V模型评估与优化方面具有巨大潜力。

英文摘要:

AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference annotations, resulting in HVEval+, the largest holistic quality assessment dataset for AI-generated human-centric videos, which comprises 1k prompts based on a comprehensive taxonomy, 20k videos generated by 24 text-to-video (T2V) models, and extensive human annotations, including 60k mean opinion scores (MOSs) and 60k preference pairs across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), as well as 20k category-specific question-answer (Q&A) pairs. Along with the HVEval+ dataset, we further propose MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific question answering within a single model. Specifically, we introduce Mixture of Projector Experts (MoPE) and Mixture of LoRA Experts (MoLE), together with a three-stage training strategy consisting of task-aware pre-training, task-specific adaptation, and adaptive routing optimization, to effectively unify multiple tasks, resulting in superior performance on both HVEval+ and Human-AGVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the HVEval+ dataset and the MoE-Rater method in advancing AI-generated video quality assessment and further facilitating the evaluation and optimization of T2V models.

补充信息

↑