arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05104cs.LG

BnBERT-iPET:基于彩票剪枝的孟加拉语稀疏少样本语言建模

BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning

  • North South University(北南大学)

机构由 AI 辅助整理,请以论文原文为准。

Sajib Hossain, Md Kamrus Samad, Anan Ghosh, Labib Imam Chowdhury, Nabeel Mohammed

AI总结:

本研究提出BnBERT-iPET,一种基于彩票剪枝的孟加拉语稀疏少样本语言建模方法,仅保留初始模型10%边、实现90%稀疏度,在孟加拉语下游任务中性能可与多个最先进语言模型媲美。

AI中文摘要:

深度神经网络凭借其复杂结构和海量边,在自然语言处理(NLP)任务中取得了令人瞩目的成功。使用BERT等大型预训练模型实现NLP的最先进性能,成本高昂、耗时且碳足迹大,难以在计算能力有限的设备上实现,这为孟加拉语等资源受限语言训练复杂模型造成了障碍。然而,在复杂神经网络中,并非所有边的影响都相同,部分边的贡献可以忽略。剪枝有望在不牺牲可比性能的前提下,减小常规网络的内存占用、缩短不断增长的网络的训练时间并提升推理效率。本研究提出了一种针对孟加拉语的稀疏少样本语言建模方法BnBERT-iPET,实验表明,仅保留BERT等初始模型10%边的轻量少样本学习语言模型,在孟加拉语等资源受限语言的挑战性任务上,性能可与大得多的模型不相上下。该模型通过迭代模式利用训练从少量样本中学习,并借助彩票假说(Lottery Ticket Hypothesis)剪枝技术实现90%的稀疏度,在孟加拉语标准基准数据集的下游任务中,可与Bangla Electra、Indic-BERT和XLM-RoBERTa等最先进语言模型一较高下。

英文摘要:

Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges. Achieving state-of-the-art performance in natural language processing with a large pre-trained model such as BERT is expensive and time-consuming, carries a large carbon footprint, and is difficult to realize on machines with minimal computational capability. This creates a barrier to training complex models for resource-constrained languages such as Bengali. However, in a complex neural model, not all edges are equally impactful, and the contributions of some of them can be neglected. Pruning promises to reduce the memory footprint of regular networks, shorten the training time of ever-growing networks, and increase inference efficiency without sacrificing comparable performance. In this work, we introduce BnBERT-iPET, a sparse few-shot language modeling approach for Bengali, and experimentally show that a lightweight few-shot-learned language model retaining only 10% of the edges of an initial model such as BERT can perform neck and neck with much larger models on challenging tasks for a resource-constrained language such as Bengali. By learning from few shots through iterative pattern exploiting training and achieving 90% sparsity with the Lottery Ticket Hypothesis pruning technique, our pruned BnBERT-iPET model proves to be a tough competitor to state-of-the-art language models such as Bangla Electra, Indic-BERT, and XLM-RoBERTa on downstream tasks over standard benchmark datasets of the Bengali language.

补充信息

↑