arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评价达罗毗荼语的专用单语与联合多语因果模型

Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

Venkata Naga Sai Vishnu Rohit Pulipaka

arXiv 2608.07727首次发表:更新:

AI 中文总结

本文训练5个GPT-2架构模型对比达罗毗荼语的单语与多语模型,发现单语模型在情感分类等任务上优于mGPT,分词器效率也更高。

AI 中文摘要

达罗毗荼语主要包括泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语,仅占多语语言模型训练数据的一小部分,因此这些模型实际保留的各语言能力尚不明确。本文从头训练了5个GPT-2架构模型,以对比4个单语模型(分别对应泰米尔语、泰卢固语、卡纳达语、马拉雅拉姆语,各有专属的32K词汇量子词分词器)与1个跨四种语言共享64K词汇量子词分词器的多语模型。所有5个模型均在清理后的CC-100、维基百科和Samanantar数据上训练,在困惑度、每字节比特数、分词器效率以及微调结果上进行测试,并与mGPT对比。实验显示,单语模型在情感分类和命名实体识别上优于mGPT,且其分词器在所有测试语言中均比共享多语模型的分词器更高效。

英文摘要

Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so it's not clear how much per-language ability these models actually keep. I have trained five GPT-2 architecture models from scratch to compare four monolingual models (one each for Tamil, Telugu, Kannada, and Malayalam, each with its own 32K-vocabulary subword tokenizer) against one multilingual model sharing a 64K-vocabulary subword tokenizer across all four languages. All the 5 models are trained on cleaned CC-100, Wikipedia, and Samanantar data. I have tested the models on perplexity, bits-per-byte, tokenizer efficiency, and fine-tuning results which are compared against mGPT. The monolingual models outperform mGPT on sentiment classification and named entity recognition, and their tokenizers proved more efficient than the shared multilingual model across all the languages tested.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑