arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04401cs.CL

低资源语言中的意识形态立场检测:孟加拉国公立与私立大学社交媒体话语中的极化现象

Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media

Safaruzzaman Shovo, Monowar Islam, Asif Hossain, Sameya Akhter, Md. Shamsul Islam

首次发表
浏览论文内容

中文总结 AI 辅助

本文构建了4,060条孟加拉语评论的标注数据集,评估多种模型检测公立与私立大学话语中的极化立场,发现零样本Llama 4 Maverick以93.31%准确率超越所有监督模型,并揭示了双方评论的不同关注点。

中文摘要 AI 辅助

公立与私立大学之争是一个备受争议的话题,并在孟加拉国的社交媒体上引发了极化现象。围绕教育质量、就业和声望的辩论在学生、家长和毕业生中十分激烈,而他们中的大多数人使用孟加拉语——一种低资源语言。为了衡量这种极化现象,本文引入了一个人工标注的数据集,包含4,060条孟加拉语评论,标注为支持公立(Pro-Public)、支持私立(Pro-Private)或中立(Neutral)。我们通过Fleiss's Kappa一致性系数评估了标注质量,该系数为0.89,表明标注者之间具有高度一致性。我们评估了经典机器学习模型(SVM、随机森林、XGBoost)、BiLSTM网络、混合BanglaBERT+XGBoost模型以及最先进的零样本大语言模型(Claude Sonnet 4、DeepSeek-V3.1、Llama 4 Maverick、Kimi K2 Thinking、Qwen3-235B Thinking)。BanglaBERT+XGBoost的准确率为91.81%,宏F1分数为91.70%,高于所有监督基线模型。零样本Llama 4 Maverick Thinking在不进行任何微调的情况下达到了0.931的宏F1分数(总体准确率为93.31%)。所有机器学习(ML)、深度学习(DL)和Transformer模型均被零样本Llama 4 Maverick模型超越。极化现象同样明显,在某些方面,支持私立的评论更为突出,这些评论强调现代化设施和按时毕业,而支持公立的评论则强调可负担性和政府工作。我们的发现为分析低资源语言中的社交媒体极化现象开辟了新方向。

英文摘要

Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduces a manually annotated dataset of 4,060 Bengali comments labeled as Pro-Public, Pro-Private, or Neutral. We evaluated the quality of our annotations by Fleiss's Kappa agreement that was 0.89, corresponding to a high agreement among annotators. The classical ML (SVM, Random Forest, XGBoost), BiLSTM network, hybrid BanglaBERT+XGBoost models and the state-of-the-art zero-shot LLMs (Claude Sonnet 4, DeepSeek-V3.1, Llama 4 Maverick, Kimi K2 Thinking, Qwen3-235B Thinking) models are evaluated. The accuracy of BanglaBERT+XGBoost is 91.81% and macro F1 score is 91.70%, which is higher than all the supervised baselines. The zero-shot Llama 4 Maverick Thinking achieves a macro F1 of 0.931 (overall accuracy of 93.31%) without any fine-tuning. All machine learning (ML), deep machine learning (DL) and transformer models were outperformed by the zero-shot Llama 4 Maverick model. Polarization also is evident, in some ways more clearly in the Pro-Private comments, which emphasize modern facilities and timely graduation, versus the Pro-Public comments, which emphasize affordability and government jobs. Our findings open new directions for analyzing social media polarization in low-resource languages.

发表机构

  • Faridpur Engineering College(法里德普尔工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑