arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17182cs.CV

BanClickThumb:用于孟加拉语YouTube视频中标题党检测的多模态数据集和Transformer融合基准测试

BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos

Md. Ariful Islam, Md Tanvirul Islam, Md. Maruf Hossain Miru, Md Khalid Syfullah

首次发表
浏览论文内容

中文总结 AI 辅助

针对孟加拉语YouTube视频标题党检测,公开多模态数据集BanClickThumb,用其对单模态和多模态方法进行基准测试,提出多模态模型BanClickFusionFormer,结合ViT和XLM - RoBERTa,证明多模态融合有效性,提供基准支持低资源多模态内容分析研究。

中文摘要 AI 辅助

标题党视频标题和缩略图夸大或歪曲内容,降低用户信任、浪费注意力并传播错误信息。检测孟加拉语标题党具有挑战性,因为公开可用的多模态数据集有限。为此,我们引入BanClickThumb,这是一个由7147个来自五个内容领域的孟加拉语YouTube缩略图-标题对组成的数据集,由十名注释者进行注释,一致性较高(科恩卡帕系数:0.83 - 0.93)。我们使用该数据集对仅文本、仅图像和多模态方法进行基准测试。单模态模型中,BanClickTextFormer(XLM - RoBERTa)准确率达0.82,BanClickImageFormer(SwiftFormer)为0.68。我们提出的多模态模型BanClickFusionFormer通过中间融合结合了ViT和XLM - RoBERTa,准确率最高达0.84。错误分析表明,密集的缩略图文本、比喻性语言和特定文化俚语仍具挑战性。我们的研究结果证明了多模态融合对孟加拉语标题党检测的有效性,并提供了一个公开可用的基准来支持未来对低资源多模态内容分析的研究。

英文摘要

Clickbait, where video titles and thumbnails exaggerate or misrepresent content, reduces user trust, wastes attention, and promotes misinformation on video-sharing platforms. Detecting Bengali clickbait remains challenging because publicly available multimodal datasets are limited. To address this gap, we introduce BanClickThumb, a curated dataset of 7,147 Bengali YouTube thumbnail-title pairs from five content domains, annotated by ten annotators with high agreement (Cohen's Kappa: 0.83-0.93). Using this dataset, we benchmark text-only, image-only, and multimodal approaches. Among unimodal models, BanClickTextFormer (XLM-RoBERTa) achieves 0.82 accuracy, while BanClickImageFormer (SwiftFormer) reaches 0.68. Our proposed multimodal model, BanClickFusionFormer, combines ViT and XLM-RoBERTa through intermediate fusion and achieves the best accuracy of 0.84. Error analysis shows that dense thumbnail text, figurative language, and culturally specific slang remain challenging. Our findings demonstrate the effectiveness of multimodal fusion for Bengali clickbait detection and provide a publicly available benchmark to support future research on low-resource multimodal content analysis.

发表机构

  • Department of Computer Science and Engineering, Bangladesh Army University of Science and Technology(孟加拉国陆军科技大学计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

↑