发表机构
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); Shenzhen University; Institute of Automation, Chinese Academy of Sciences; Pengcheng Laboratory; Tongji University; Tsinghua University; The Chinese University of Hong Kong; University of Trento; Huawei(广东省人工智能与数字经济实验室(深圳); 深圳大学; 中国科学院自动化研究所; 鹏城实验室; 同济大学; 清华大学; 香港中文大学; 特伦托大学; 华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了HUG-VIS多模态基准,包含30名演员的8400段同步多模态视频等数据,评估了四类任务的模型,揭示了情感识别、生成等任务的关键特性与现存问题。
AI 中文摘要
视觉智能旨在感知、解释和合成视觉世界,是现代计算机视觉的核心。以人为中心的视觉智能要求极高,因为它将人视为具有表现力、处于社会情境中的主体,其意义很少仅通过外观传达。它将视觉与音频、语言相结合,覆盖四个代表性任务:人类情感识别、人类视频生成、人类语音克隆和人类视频抠像。然而,现有资源仍针对特定任务,为单个问题提供模态和标注,而非协调理解与生成的共享基础,这限制了多模态信号的使用及更广泛的研究。我们通过HUG-VIS解决这一缺口,它是面向视觉智能中以人为中心的理解与生成的统一基准。该基准包含30名专业演员的8400段坐姿半身视频,每位演员在受控的普通话工作室协议下执行相同的280个情感-动作-提示任务,配有同步视频、音频、文本和Alpha抠像。我们在统一的零样本协议下,使用自动指标、特定标准的平均意见得分及多项跨任务分析,评估了四个任务中多种开源和闭源模型。结果显示:(i)语言内容主导当前的情感识别,而纯视觉情感识别表现最弱;(ii)在视频生成和语音克隆中,自动指标与人类判断总体一致,但在排名前列存在差异,需联合报告;(iii)运动下的边界保真度是人类抠像的主要剩余障碍;(iv)任务难度随情感、模型和指标变化,存在显著的跨任务相关性。数据集和结果可在该httpsURL获取。
英文摘要
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task-specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open- and closed-source models across the four tasks under a unified zero-shot protocol using automatic metrics, criterion-specific mean opinion scores, and multiple cross-task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross-task correlations. The dataset and results are available at https://github.com/GML-MMGroup/HUG-VIS.