arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HelaBERT:采用双池化分类头增强僧伽罗语理解能力

HelaBERT: Enhancing Sinhala Language Understanding with Dual Pooling Classification Head

Thisen Ekanayake, Nisansa de Silva

arXiv 2608.22922首次发表:更新:

发表机构

University of Moratuwa(莫拉图瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出两款僧伽罗语BERT模型HelaBERT-Small与HelaBERT-Large,采用定制分词器,在四项下游分类任务上评估,提出双池化分类头并发布模型以推进僧伽罗语NLP研究。

AI 中文摘要

我们提出了HelaBERT,这是两款基于BERT的掩码语言模型,从头开始在约10亿个僧伽罗语文本令牌上进行预训练,这些文本来自MADLAD-400、CulturaX以及包含新闻文章、僧伽罗语维基百科和网络爬取数据的自定义语料库。HelaBERT-Small(约2330万个参数,6层)和HelaBERT-Large(约1.1亿个参数,12层)均使用针对僧伽罗语黏着形态和复杂脚本定制的SentencePiece Unigram分词器(词汇量为32000)。我们在四项下游僧伽罗语文本分类任务上评估这两款模型:新闻类别分类、新闻来源分类、情感分析和写作风格分类,采用5次独立随机种子运行,使用分层80/20训练/测试分割。我们还提出了一种双池化分类头,并在所有四项任务上对其进行系统评估,发现其在情感分析上持续改进,在新闻类别分类上为HelaBERT-Small带来适度提升,而标准的<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>-线性分类头在新闻来源分类(一项平均输入长度较短的标题级任务)上仍具有竞争力。我们发布这两款模型以支持僧伽罗语自然语言处理的进一步研究。

英文摘要

We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX, and a custom corpus comprising news articles, Sinhala Wikipedia, and web crawl data. HelaBERT-Small (~23.3M parameters, 6 layers) and HelaBERT-Large (~110M parameters, 12 layers) both use a SentencePiece Unigram tokenizer (vocabulary size 32,000) tailored to Sinhala's agglutinative morphology and complex script. We evaluate both models on four downstream Sinhala text classification tasks: news category classification, news source classification, sentiment analysis, and writing style classification, using 5 independent seed runs with stratified 80/20 train/test splits. We additionally propose a dual pooling classification head and evaluate it systematically across all four tasks, finding consistent improvements on sentiment analysis and a moderate gain on news category classification for HelaBERT-Small, while the standard [CLS]-linear head remains competitive on news source classification, a headline-level task with short average input length. We release both models to support further research in Sinhala NLP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑