ConlangBench:通过多样人造语言探索大语言模型的语言知识与学习
ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages
AI总结:
研究人员提出首个针对21种现有人造语言的大规模基准ConlangBench,收集大量平行句对和词汇条目,通过双向翻译实验发现模型在后天人造语言上表现更好,且可学习8种拥有足够语料库的人造语言,为人造语言探索LLMs低资源语言学习提供独特试验台。
AI中文摘要:
人造语言(conlangs)是人为创造的人类语言,拥有丰富的语言创造力传统。尽管它们在研究大语言模型(LLMs)的语言学习方面具有潜力,但现有人造语言在LLM研究中仍未得到充分探索。我们提出ConlangBench,这是首个针对21种现有人造语言评估和训练LLMs的大规模基准。我们收集了超过2100万个人造语言-英语平行句对(包括20种非世界语人造语言的43万对)以及32.1万个词汇条目。在双向翻译实验中,我们发现模型在后天人造语言(其词汇源自自然语言)上表现更好,这反映了人造语言的设计特征。在ConlangBench上训练还表明,模型可以学习所有8种拥有足够平行语料库的人造语言,而它们的学习曲线取决于人造语言的创造方式。我们的发现表明,人造语言为研究LLMs如何获取低资源语言提供了独特的试验台。
英文摘要:
Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for studying language learning in large language models (LLMs), existing conlangs remain largely underexplored in LLM research. We present ConlangBench, the first large-scale benchmark for evaluating and training LLMs on 21 existing conlangs. We collect over 21M conlang-English parallel sentence pairs (including 430K pairs across the 20 non-Esperanto conlangs) and 321K vocabulary entries. In bidirectional translation experiments, we find that models perform better on a posteriori conlangs, whose vocabularies are derived from natural languages, reflecting the design characteristics of conlangs. Training on ConlangBench also shows that models can learn all eight conlangs for which sufficient parallel corpora are available, while their learning curves vary depending on how the conlangs were created. Our findings suggest that conlangs provide a unique testbed for investigating how LLMs acquire low-resource languages.