arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MultiHuSE:一个用于幽默风格与情绪的多模态数据集

MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat

arXiv 2609.11322首次发表:更新:

发表机构

Imperial College London; California Polytechnic State University(伦敦帝国理工学院; 加州州立理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MultiHuSE是一个包含2,407个视频的多模态数据集,覆盖四种幽默风格和情绪标注,通过多模态融合在幽默分类中达到80.1%的准确率,为幽默与情绪研究提供实证支持。

AI 中文摘要

言语幽默的计算识别仍然是一项具有挑战性的任务,需要理解语言、表达风格、情绪和文化背景。现有的大多数方法侧重于二元分类,并且缺乏能够捕捉幽默心理维度以及表达变化的数据集。我们引入了MultiHuSE,这是一个多模态数据集,包含50位人口统计学多样化的演员表演1,463个文本样本的2,407个高清视频,涵盖四种心理幽默风格(亲和型、攻击型、自我提升型和自嘲型)以及中性内容。其中一部分子集还额外标注了潜在情绪。该数据集独特地捕捉了同一文本的多种演员诠释,从而能够系统分析表达多样性。基线实验表明,在幽默风格分类中,多模态融合优于单模态方法(准确率80.1%对77.4%),其中对亲和型幽默的提升尤为显著(从66%提升到74%)。虽然文本提供了最强的单一信号,但融合模型带来了有意义的改进。我们希望MultiHuSE能够为将幽默与情绪联系起来的心理学理论提供实证支持,同时也为人类交流、福祉和人工智能驱动的交互研究开辟新途径。该数据集在最终用户许可协议下可供学术使用。

英文摘要

Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions of humour alongside variations in expression. We introduce MultiHuSE, a multimodal dataset comprising 2,407 high-definition videos of 50 demographically diverse actors performing 1,463 text samples across four psychological humour styles (affiliative, aggressive, self-enhancing, and self-deprecating), as well as neutral content. A subset is additionally annotated for underlying emotions. The dataset uniquely captures multiple actor interpretations of the same texts, enabling systematic analysis of expressive diversity. Baseline experiments show that multimodal fusion outperforms unimodal approaches (80.1% vs. 77.4% accuracy) in humour style classification, with particularly strong gains for affiliative humour (66% to 74%). While text provides the strongest individual signal, fusion models deliver meaningful improvements. We hope that MultiHuSE provides empirical support for psychological theories linking humour and emotion, while also opening new avenues for research in human communication, well-being, and AI-driven interaction. The dataset is available for academic use under an End-User Licence Agreement.

Comments7 pages, 3 figures, 5 tables. Accepted at IEEE CBMI 2025 (International Conference on Content-Based Multimedia Indexing), Dublin, Ireland

Journal ref2025 International Conference on Content-Based Multimedia Indexing (CBMI), Dublin, Ireland, 2025, pp. 1-7

DOI:10.1109/CBMI66578.2025.11339313

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑