arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20984cs.CVcs.CLcs.CY

MigrationNarrate:用于检测YouTube视频中移民叙事的数据集

MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos

  • University of Sheffield(谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

Fatima Haouari, Carolina Scarton, Kalina Bontcheva

AI总结:

该研究推出首个用于检测英国移民叙事的多模态数据集MigrationNarrate,包含1115个YouTube视频字幕,结合预训练编码器与大语言模型开展基准测试,为移民叙事检测研究提供关键资源。

AI中文摘要:

叙事是社会沟通框架的核心,因此检测叙事对于理解和分析公共话语至关重要。现有研究已在不同领域探索了叙事的检测与提取,但移民叙事的研究仍严重不足,主要原因是缺乏专门的带注释数据集。此外,公共沟通最近转向以视频为中心的平台,叙事通过多模态信号传递并被大规模消费;尽管有这一转变,视频中的叙事仍在很大程度上未被探索。为弥合这些差距,我们推出MigrationNarrate,这是首个用于检测英国移民叙事的多模态数据集,包含1115个YouTube视频的字幕,采用由12个移民超级叙事和53个叙事标签组成的两级分类法进行注释。本文详细介绍了该数据集的设计、收集与注释过程,以及使用预训练编码器模型结合开源和闭源大型语言模型得到的基准结果,最后通过全面的错误分析为未来工作提供见解。

英文摘要:

Narratives are central to how social communication is framed, making their detection critical for understanding and analysing public discourse. Prior work has explored narrative detection and extraction across diverse domains; however, migration narratives remain significantly understudied, primarily due to the absence of dedicated annotated datasets. Furthermore, public communication has recently shifted towards video-centric platforms, where narratives are conveyed through multimodal signals and consumed at scale. Despite this shift, narratives in videos remain largely unexplored. To bridge these gaps, we introduce MigrationNarrate, the first multimodal dataset for detection of migration narratives in the UK, consisting of 1,115 YouTube video transcripts annotated using a two-level taxonomy of 12 migration super-narratives and 53 narrative labels. This paper details the dataset design, collection, and annotations; together with benchmark results using a combination of pre-trained encoder models and both open- and closed-source Large Language Models. Finally, a thorough error analysis offers insights for future work.

补充信息

↑