Video2Reaction:训练基础视频模型以预测观众反应
Video2Reaction: Training Foundation Video Models to Predict Audience Reaction
浏览论文内容
中文总结 AI 辅助
本研究构建多模态数据集Video2Reaction,对采用LoRA微调的VLMs进行基准测试,发现其可有效学习并迁移至VCE数据集,仅用1%VCE训练数据适配的LLaVA-NeXT-Video-7B可达到与全数据训练相当的性能。
中文摘要 AI 辅助
我们提出Video2Reaction,这是一个多模态数据集,将短视频片段与野外环境中观众通过社交媒体评论表达的情绪反应建立映射。Video2Reaction通过大规模聚合在线评论中的反应,捕捉情绪反应的自然多样性,将标签建模为分类情绪的分布,以更好地反映情绪感知的主观性和模糊性。我们对两个采用LoRA微调的视觉语言模型(VLMs)进行基准测试,结果显示VLMs可从Video2Reaction中有效学习,在主导反应预测任务上优于专门的基线模型。我们进一步证明,在Video2Reaction上预微调的VLMs可有效迁移至VCE——另一个具有不同分类体系和视频领域的诱导情绪数据集。值得注意的是,在Video2Reaction上预微调且仅用VCE训练数据的1%进行适配的LLaVA-NeXT-Video-7B,达到了0.682的Top-3准确率,与在完整VCE数据集上训练得到的最佳报告性能相当。该数据集可通过此URL获取。
英文摘要
We introduce Video2Reaction, a multimodal dataset that maps short movie segments to the induced emotional reactions of viewers in the wild, as expressed through social media comments. Video2Reaction captures the natural diversity of emotional responses by aggregating reactions from online comments at scale, modeling labels as distributions over categorical emotions to better reflect the subjective and ambiguous nature of emotional perception. We benchmark two vision-language models (VLMs) finetuned with LoRA, showing that VLMs learn effectively from Video2Reaction and outperform specialized baselines on dominant reaction prediction. We further demonstrate that VLMs pre-finetuned on Video2Reaction transfer effectively to VCE, another induced emotion dataset with a different taxonomy and video domain. Notably, LLaVA-NeXT-Video-7B pre-finetuned on Video2Reaction and adapted on only 1% of VCE training data achieves a top-3 accuracy of 0.682, on par with the best reported VCE performance trained on the full dataset. The dataset is available at https://huggingface.co/datasets/infofusionlab/Video2Reaction
发表机构
- UMass Amherst(马萨诸塞大学阿默斯特分校)
- Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。