arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25190cs.CLcs.LG

BanglaMamba:探索用于孟加拉语假新闻检测的状态空间模型

BanglaMamba: Exploring State Space Models for Bangla Fake News Detection

  • BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

M. K. Khalidi Siam

中文总结 AI 辅助

本文提出BanglaMamba,将基于Mamba的状态空间模型用于孟加拉语假新闻检测,其性能与从头训练的CustomBERT相当,且推理效率优于BERT模型,凸显了该模型在资源受限场景的应用潜力。

中文摘要 AI 辅助

由于错误信息通过在线新闻平台和社交媒体快速传播,假新闻检测已成为一项重要的自然语言处理(NLP)任务。尽管诸如BanglaBERT之类的基于Transformer的模型在孟加拉语文本分类中表现出强劲性能,但其二次计算复杂度使其不太适合资源受限环境下的长文档处理。本文研究基于Mamba的状态空间模型(SSMs)作为孟加拉语假新闻检测的高效替代方案,提出了BanglaMamba并将其与预训练的BanglaBERT以及从头开始训练的同等配置BERT模型进行比较。实验结果显示,BanglaBERT的Macro-F1分数最高,为0.9260;BanglaMamba的Macro-F1分数为0.9029,尽管架构不同,但其性能与从头训练的CustomBERT(0.9057)相当。同时,BanglaMamba的推理吞吐量比基于BERT的模型高约2.2倍,推理峰值GPU内存使用率低49%。跨数据集评估显示,BanglaBERT对外部数据集的泛化能力更强,凸显了大规模预训练的重要性。这些发现表明,基于Mamba的SSMs可作为基于Transformer架构的具有竞争力且计算高效的替代方案,用于孟加拉语假新闻检测,尤其适用于资源受限的场景。

英文摘要

Fake news detection has become an important Natural Language Processing (NLP) task due to the rapid spread of misinformation through online news platforms and social media. While transformer-based models such as BanglaBERT achieve strong performance for Bangla text classification, their quadratic computational complexity makes them less suitable for long-document processing in resource-constrained environments. This paper investigates Mamba-based State Space Models (SSMs) as an efficient alternative for Bangla fake news detection. We propose BanglaMamba and compare it with pre-trained BanglaBERT and a similarly configured BERT model trained from scratch. Experimental results show that BanglaBERT achieves the highest Macro-F1 score (0.9260), while BanglaMamba (0.9029) achieves performance comparable to the from-scratch CustomBERT (0.9057) despite using a different architecture. Meanwhile, BanglaMamba achieves approximately $2.2\times$ higher inference throughput and 49% lower inference peak GPU memory usage than the BERT-based models. Cross-dataset evaluation shows that BanglaBERT generalizes better to an external dataset, highlighting the importance of large-scale pretraining. These findings demonstrate that Mamba-based SSMs can provide a competitive and computationally efficient alternative to Transformer-based architectures for Bangla fake news detection, particularly in resource-constrained settings.

↑