arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLMs)从针对性的合成多语言数据中提升性能

LLMs Get Smarter from Targeted Synthetic Multilingual Data

Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen, Neha Gupta, Andreas Stolcke

arXiv 2608.15964首次发表:更新:

发表机构

UIUC; Uniphore(伊利诺伊大学厄巴纳-香槟分校; 优尼佛(Uniphore)公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出HOTFIXR数据生成框架,通过探测学生模型多语言弱点生成合成多语言数据,可提升大语言模型的多语言性能,降低微调引发的灾难性遗忘。

AI 中文摘要

语言特定能力(LSC)是指语言模型的性能会因提示语的语言不同而出现差异的现象,换句话说,当用不同语言提示同一语义查询时,语言模型会输出不同(甚至可能错误)的响应。现有研究将此归因于模型跨语言语义表示的内部不对齐。目前,文献中解决LSC的主要方法有两种:(1)将所有查询路由至英语,可提升性能,但会将语言表达性限制在英语范围内;(2)使用语言均衡的数据进行训练,可均衡模型在各语言上的性能,但会降低整体性能。本研究从数据中心视角出发,提出HOTFIXR(Hardness Optimized Training data For Improving X-Lingual Reasoning,用于提升跨语言推理的难度优化训练数据),这是一种数据生成框架,它利用模型探测并学习学生模型的多语言弱点,进而生成数据以缓解这些弱点。HOTFIXR可生成合成多语言训练数据,从而提升多语言性能。我们在3个分布内任务、3个分布外任务以及4种分布外语言上进行评估,结果显示,平均而言,HOTFIXR(1)将分布内性能提升6.2%;(2)将微调引发的分布外(OOD)任务上的灾难性遗忘降低3.7%;(3)在分布外语言上提升7.1%。总体而言,由于许多实际应用需要多语言大语言模型,本研究为提升大语言模型的多语言熟练度作出了贡献,我们将在论文接收后发布代码。

英文摘要

Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic query when prompted in different languages. Prior work attributes this to an internal misalignment of semantic representation across languages. Currently, there are two main approaches to address LSC in the literature: (1) routing all queries through English, improving performance, but limiting language expressivity to English; or (2) training on language-balanced data, equalizing model performance across languages, but reducing overall performance. In this work, we take a data centric perspective and introduce HOTFIXR: Hardness Optimized Training data For Improving X-Lingual Reasoning. It is a data generation framework that uses models to probe and learn a student model's multilingual weaknesses, and generates data to mitigate them. HOTFIXR can generate multilingual synthetic training data that can improve multilingual performance. We evaluate on three in-distribution tasks, three out-of-distribution tasks, and four out-of-distribution languages. On average, HOTFIXR (1) improves in-distribution performance by 6.2%, (2) reduces catastrophic forgetting (induced by fine-tuning) on OOD tasks by 3.7%, and (3) on OOD languages by 7.1%. Overall, as many real-world applications requires multilingual LLMs, our work contributes to the efforts of making LLMs multilingually proficient. We will release code upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑