arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16010cs.CL

通过生成式与抽取式预训练Transformer实现尼泊尔法律专业知识(NepLEGiT)

Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT)

Ranjit Raut, Tishya Dhakal, Aaryan Shakya, Bhabuk Thapa, Prasiddha Koirala, Bal Krishna Bal

首次发表
浏览论文内容

中文总结 AI 辅助

针对尼泊尔法律信息获取难的问题,提出NepLEGiT小型语言模型,基于约400万token法律文本预训练GPT-2,在验证集上困惑度1.8,准确率82.9%,并对比mBERT和MuRIL编码器基线,以普及法律知识。

中文摘要 AI 辅助

法律语言的复杂性和法律信息获取渠道的有限性,对尼泊尔的司法公正构成了重大挑战。由于语言障碍、信息碎片化以及法律专业知识的严重短缺(尤其是在农村地区),许多公民仍然无法获得传统的法律服务。我们提出了NepLEGiT(Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers),这是一个专门设计的小型语言模型(SLM),旨在普及法律知识并提升尼泊尔的法律服务交付。我们从零开始,在一个包含约400万token的尼泊尔法律文本精选语料库上预训练了一个基于解码器的GPT-2 SLM,该语料库涵盖宪法、民法和刑法以及行政法规。该模型包含约3000万个参数,采用6层、6头、384维的Transformer架构,并使用warmup余弦衰减调度、梯度累积和混合精度运算进行训练。在留出的验证集上,NepLEGiT达到了0.5684的交叉熵损失、1.8的困惑度和82.9%的下一个词元预测准确率。我们还进一步评估了在同一语料库上对mBERT和MuRIL进行持续掩码语言模型预训练的效果;mBERT达到了2.35的困惑度(评估损失0.8565),优于MuRIL(困惑度6.07,评估损失1.8026),提供了一个与NepLEGiT的生成式定位互补的强大编码器基线。

英文摘要

The complexity of legal language and limited accessibility to legal information pose significant challenges to justice delivery in Nepal. Traditional legal services remain inaccessible to many citizens due to language barriers, information fragmentation, and a critical shortage of legal expertise, particularly in rural areas. We present NepLEGiT (Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers), a specialized small language model (SLM) designed to democratize legal knowledge and enhance legal-service delivery in Nepal. We pre-train a decoder-based GPT-2 SLM from scratch on a curated corpus of ~4 million tokens of Nepali legal text, covering constitutional law, civil and criminal codes, and administrative regulations. The model comprises ~30 million parameters in a 6-layer, 6-head, 384-dimensional transformer trained with warmup cosine-decay scheduling, gradient accumulation, and mixed-precision arithmetic. On a held-out validation split, NepLEGiT attains a cross-entropy loss of 0.5684, a perplexity of 1.8, and a next-token prediction accuracy of 82.9%. We further evaluate continual masked-language-model pre-training of mBERT and MuRIL on the same corpus; mBERT achieves a perplexity of 2.35 (eval loss 0.8565), outperforming MuRIL (perplexity 6.07, eval loss 1.8026), providing a strong encoder baseline complementary to NepLEGiT's generative orientation.

发表机构

  • Kathmandu University(加德满都大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑