arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BanglaDial-Abuse:用于孟加拉语滥用文本中区域方言识别的语料库基础数据集

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

Hasin Almas Sifat

arXiv 2610.01150首次发表:更新:

发表机构

American International University-Bangladesh(美国国际大学孟加拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对孟加拉语滥用文本的区域方言识别挑战,构建了包含四种方言各250句的平衡数据集,并采用基于语料库的合成方法生成,主要贡献在于提供方言识别任务资源。

AI 中文摘要

区域语言变异仍然是孟加拉语自然语言处理中的一个重要挑战,尤其是在非正式和非标准文本中。本文介绍了BanglaDial-Abuse,一个平衡的孟加拉文脚本数据集,用于识别滥用和敌意孟加拉语文本中的区域方言。该数据集包含1,000个句子,均匀分布在四种语言变体中:标准孟加拉语、Chattagram、Sylhet和Barishal,每个类别有250个样本。该资源采用基于语料库的合成程序构建,融合了代词、所有格形式、动词形态、否定、疑问结构、后置词、词汇和孟加拉文拼写惯例的区域变异,同时保留了潜在的敌意或滥用含义。描述性分析显示,四个类别的句子长度分布大致相当,但词汇空间部分不同。两两Jaccard词汇相似度范围从0.37到0.56。主要任务是四类区域方言识别,而非二元的滥用文本检测。该数据集通过Zenodo在知识共享署名4.0许可下公开可用。当前版本旨在作为研究和原型制作语料库,而非母语者验证的金标准语言资源。关键词:孟加拉语、孟加拉语、方言识别、区域方言、滥用语言、低资源NLP、Chattagram、Sylhet、Barishal、数据集

英文摘要

Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class. The resource was constructed using a corpus-grounded synthetic procedure incorporating regional variation in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, vocabulary, and Bengali-script spelling conventions while preserving the underlying hostile or abusive meaning. Descriptive analysis shows broadly comparable sentence-length distributions but partially distinct lexical spaces across the four classes. Pairwise Jaccard vocabulary similarity ranges from 0.37 to 0.56. The primary task is four-class regional dialect identification rather than binary abusive-text detection. The dataset is publicly available through Zenodo under a Creative Commons Attribution 4.0 license. The current version is intended as a research and prototyping corpus rather than a native-speaker-validated gold-standard linguistic resource. Keywords: Bangla, Bengali, dialect identification, regional dialect, abusive language, low-resource NLP, Chattagram, Sylhet, Barishal, dataset

Comments5 pages, 3 figures, 1 table. Dataset Version 1.0 available on Zenodo: 10.5281/zenodo.23074319

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑