arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29045cs.CV

基于情感异质图推理与多任务联合学习的自适应情感视频描述

Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning

Junbo Wang, Liangyu Fu, Yuke Li, Xuecheng Wu, Zhiyong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有情感视频描述方法情感先验处理的缺陷,提出基于情感异质图与多任务联合学习的SAGML框架,实现了更鲁棒的情感视频描述性能。

中文摘要 AI 辅助

情感视频描述(EVC)旨在以事实准确性与情感表现力描述视频,要求模型感知微妙、模糊且随时间变化的情感线索,并将其转化为自然语言,同时不削弱客观视觉内容。现有方法逐步引入了上下文注意力、情感解释、情感先验、动态情感感知及情感-原因推理,但多数仍依赖全局情感向量或刚性层次先验。近期方法中,树状情感先验建立了心理情感类别与日常情感词汇间由粗到细的联系,但其硬从属掩码一旦粗类别预测不准确,便会不可逆地抑制正确的词汇情感,且难以表示真实视频中频繁出现的混合或重叠情感。为解决上述问题,本文提出SAGML,一种基于情感异质图与多任务语言建模的自适应EVC框架。SAGML未将情感先验视为离散树,而是构建了包含目录级情感节点与词汇级情感词汇节点的软情感异质图,将软门作为连续偏置注入视频到情感的图注意力中,使视觉支持的词汇情感可恢复,而非被硬掩码去除。生成的情感表示与视觉标记一同输入因果语言解码器,同时双目录与词汇头对提示隐藏状态施加显式情感分布学习。整体模型采用自回归描述生成与情感分布监督相结合的联合目标训练,SAGML为EVC提供了抗错误且多情感感知的基线。

英文摘要

Emotional video captioning (EVC) aims to describe a video with both factual correctness and affective expressiveness. It requires a model to perceive subtle, ambiguous, and temporally varying emotional cues and translate them into natural language without weakening objective visual content. Existing methods have progressively introduced contextual attention, emotion interpretation, emotion priors, dynamic emotion perception and emotion-cause reasoning. Nevertheless, most of them still depend on either global emotion vectors or rigid hierarchical priors. In recent methods, the tree-structured emotion prior establishes a coarse-to-fine connection between psychological emotion categories and daily emotion words, but its hard subordinate masking may irreversibly suppress correct lexical emotions once the coarse category prediction is inaccurate. It is also limited in representing mixed or overlapping emotions that frequently occur in real videos. To address the issues, we propose SAGML, an adaptive EVC framework via affective heterogeneous graph and multi-task language modeling. Instead of treating the emotion prior as a discrete tree, SAGML constructs a soft affective heterogeneous graph containing catalog-level emotion nodes and lexical-level emotion word nodes. The soft gate is injected into video-to-emotion graph attention as a continuous bias, allowing visually supported lexical emotions to remain recoverable rather than being removed by a hard mask. The resulting affective representation is fed together with visual tokens into a causal language decoder, while dual catalog and lexical heads impose explicit emotion distribution learning on the prompt hidden states. The overall model is trained with a joint objective that combines autoregressive caption generation and emotion distribution supervision. SAGML provides an error-resilient and multi-emotion-aware baseline for EVC.

发表机构

  • School of Software, Northwestern Polytechnical University(西北工业大学软件学院)
  • School of Computer Science, The University of Sydney(悉尼大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

↑