面向高效语言模型推理的跨语言词表适配实证研究
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
- School of Computer Science, University of Sheffield(谢菲尔德大学计算机学院)
- The Alan Turing Institute(阿兰·图灵研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文对五种跨语言词表适配方法在四种生成式LLMs和四种语言上进行实证研究,发现CVA能实现高达271.5%的推理加速,且在均衡多语言数据上预训练的模型适配后下游性能与原模型相当。
AI中文摘要:
最先进的生成式大型语言模型(LLMs)的发展不成比例地依赖于以英语为中心的分词器、词表和预训练数据。尽管部分 LLMs 具备多语言能力,但近期研究表明,生成英语以外的语言文本时,其推理效率会下降。这导致推理时间和成本增加。跨语言词表适配(CVA)方法被提出用于将模型适配至目标语言,旨在提升下游性能。然而,这些方法在提升生成式 LLMs 推理效率方面的有效性仍有待探索。本文对五种 CVA 方法在四个生成式 LLMs(包括单语和多语言模型)上进行了实证研究,涵盖四种类型学上不同的语言和四个自然语言理解任务。我们发现,CVA 显著促进了 LLM 推理加速,最高达 271.5%。我们还表明,适配在更均衡的多语言数据上预训练的 LLMs,可使其下游性能与原始模型相当。
英文摘要:
The development of state-of-the-art generative large language models (LLMs) disproportionately relies on English-centric tokenizers, vocabulary and pre-training data. Despite the fact that some LLMs have multilingual capabilities, recent studies have shown that their inference efficiency deteriorates when generating text in languages other than English. This results in increased inference time and costs. Cross-lingual vocabulary adaptation (CVA) methods have been proposed for adapting models to a target language aiming to improve downstream performance. However, the effectiveness of these methods on increasing inference efficiency of generative LLMs has yet to be explored. In this paper, we perform an empirical study of five CVA methods on four generative LLMs (including monolingual and multilingual models) across four typologically-diverse languages and four natural language understanding tasks. We find that CVA substantially contributes to LLM inference speedups of up to 271.5\%. We also show that adapting LLMs that have been pre-trained on more balanced multilingual data results in downstream performance comparable to the original models.