arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09464cs.CL

BanglaRhet:孟加拉语政治演讲中修辞与说服检测的经典模型与Transformer模型基准评测

BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech

发表机构孟加拉国东北大学
查看机构详情
  • North East University Bangladesh(孟加拉国东北大学)

机构由 AI 辅助整理,请以论文原文为准。

Rohit Kumar Sen, Anik Chowdhury

首次发表
浏览论文内容

中文总结 AI 辅助

本文构建BanglaRhet语料库(30,289个孟加拉语政治演讲片段),基准评测四种Transformer模型,发现BanglaBERT在修辞与说服检测中宏F1分别达65.40%和66.46%,优于经典基线,为孟加拉语政治话语分析提供基准。

中文摘要 AI 辅助

政治话语常常使用修辞性和说服性语言来构建叙事、影响公众舆论并动员受众。尽管孟加拉语自然语言处理在情感分析和意见挖掘方面取得了进展,但对Transformer模型在孟加拉语政治演讲中进行细粒度修辞和说服技巧检测的系统性基准评测仍然在很大程度上未被探索。本文提出了一项基于Transformer模型的基准研究,用于检测孟加拉语政治话语中的修辞形式和说服意图。利用BanglaRhet——一个从公开可获取的政治新闻来源收集的包含30,289个孟加拉语政治演讲片段的人工标注语料库,我们构建了两个监督式单标签分类任务:修辞技巧检测(对比、重复、夸张、隐喻、反问)和说服技巧检测(归咎、行动号召、团结号召、道德诉求、情感诉求和逻辑诉求)。我们评估了四种基于Transformer的模型,即BanglaBERT、BanglaBERT-Base、SahajBERT和XLM-RoBERTa-Base,并与经典的TF-IDF基线进行了比较。BanglaBERT取得了最高性能,在修辞技巧检测中宏F1分数达到65.40%,在说服技巧检测中达到66.46%,分别比最佳调优的经典基线高出19.2和13.8个宏F1百分点。类别层面的分析表明,错误主要与标签之间的语义重叠、比喻性语言和类别不平衡有关。这些结果为孟加拉语修辞和说服感知的政治话语分析提供了初步的基准基线,并凸显了对上下文感知和多标签建模的需求。

英文摘要

Political discourse often uses rhetorical and persuasive language to frame narratives, influence public opinion, and mobilize audiences. While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored. This paper presents a benchmark study of transformer-based models for detecting rhetorical form and persuasive intent in Bangla political discourse. Using BanglaRhet, a manually annotated corpus of 30,289 Bangla political speech segments collected from publicly available political news sources, we formulate two supervised single-label classification tasks: rhetorical technique detection (contrast, repetition, exaggeration, metaphor, rhetorical questions) and persuasion technique detection (blame assignment, call to action, unity call, moral, emotional, and logical appeals). We evaluate four transformer-based models, BanglaBERT, BanglaBERT-Base, SahajBERT, and XLM-RoBERTa-Base, against classical TF-IDF baselines. BanglaBERT achieves the highest performance, with 65.40% macro-F1 for rhetorical technique detection and 66.46% for persuasion technique detection, outperforming the best tuned classical baseline by 19.2 and 13.8 macro-F1 points, respectively. Class-level analysis indicates that errors are mainly associated with semantic overlap among labels, figurative language, and class imbalance. The results provide initial benchmark baselines for Bangla rhetorical and persuasion-aware political discourse analysis and highlight the need for context-aware and multi-label modeling.

补充信息

↑