AI 中文总结
本文针对多语言神经机器翻译的高内存消耗问题,提出语料驱动词汇剪枝结合定向微调的框架,在英语-阿拉伯语对实验中实现60%内存节省且性能无损失,优化模型表现优于专用双语基线。
AI 中文摘要
将大型预训练多语言模型应用于神经机器翻译(MNMT)面临一项重大挑战:由于词汇表和嵌入层过大,导致内存与计算消耗过高。尽管现有的剪枝、量化、知识蒸馏等压缩方法可减少参数冗余,但它们主要保留原始词汇表结构,从而未解决主要的低效来源。本文提出一种通用优化框架,将词汇剪枝方法与针对MNMT模型的定向微调协议相结合。我们使用M2M100、NLLB-200、mBART-50三种模型,以英语-阿拉伯语语言对对所提框架进行评估。我们的方法将词汇表规模从超过128000个token缩减至约10000个,实现60%的内存节省且性能无损失。结果表明,优化后的多语言模型可达到或超过专用双语基线的性能。具体而言,经剪枝和微调的M2M100模型取得了42.04的BLEU得分(对比OPUS-MT-en-ar双语模型的44.59),同时在COMET指标上显著优于后者(0.8730对比0.7911),展现出更优的语义充分性与流畅性。
英文摘要
The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly large vocabularies and embedding layers. Although existing compression methods like pruning, quantization and knowledge distillation reduce parameter redundancy, they mainly preserve the structure of the original vocabulary, thereby leaving a major source of inefficiency unresolved. We propose in this paper a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models. We evaluate the proposed framework using three models (M2M100, NLLB-200, mBART-50) on the English-Arabic language pair. Our approach reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance. Results show that optimized multilingual models can match or exceed the performance of dedicated bilingual baselines. In particular, the pruned and fine-tuned M2M100 model achieves a competitive BLEU score of 42.04 (against 44.59 for the OPUS-MTen- ar bilingual model) while it significantly outperforms it on the COMET metric (0.8730 vs 0.7911) revealing superior semantic adequacy and fluency.