arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EuroAlpaca:面向欧洲语言的任务保留型指令数据本地化方案

EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages

Aleix Sant, Jordi Luque, Carlos Escolano

arXiv 2609.05043首次发表:更新:

发表机构

Scientific Research, Telefónica Innovación Digital; Universitat Politècnica de Catalunya(西班牙电信创新数字科研部; 加泰罗尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EuroAlpaca是面向50种欧洲语言的任务保留型指令数据本地化方案,搭配多语言基准European-IFEval,可解决直接机器翻译导致的指令微调数据损坏问题,提升模型在多语言指令遵循任务上的准确率与生成质量。

AI 中文摘要

机器翻译(MT)为将英文指令微调数据扩展至多语言提供了可扩展的方式,但它可能会扭曲任务关键约束和所需输出,从而产生损坏的训练示例并降低基于此类数据训练的模型性能。我们推出EuroAlpaca,这是一个覆盖50种欧洲语言及区域变体的任务保留型本地化流程与近并行资源,同时推出European-IFEval,这是一个用于可验证指令遵循的多语言基准。根据示例的不同,我们的流程会应用分领域机器翻译以保留任务关键内容,或重构与任务等效的目标语言实例,随后验证跨领域连贯性和目标语言一致性。在针对四个大语言模型(LLM)的LoRA实验中,使用直接翻译数据训练会提升Aya评估套件上的ROUGE-L和F-BERT分数,但会使European-IFEval的准确率相较于未适配基准降低29.8%;相比之下,使用EuroAlpaca适配则使准确率较同一基准提升12.9%,逆转了直接机器翻译造成的性能下降,同时在Aya上也取得了最高的ROUGE-L和F-BERT分数。这些结果表明,保留任务语义对于多语言指令微调至关重要。

英文摘要

Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and degrading models trained on such data. We introduce EuroAlpaca, a task-preserving localisation pipeline and near-parallel resource covering 50 European languages and regional varieties, together with European-IFEval, a multilingual benchmark for verifiable instruction following. Depending on the example, our pipeline applies field-wise MT while preserving task-critical content or reconstructs a task-equivalent target-language instance, followed by validation of cross-field coherence and target-language consistency. Across LoRA experiments with four LLMs, training on directly translated data improves ROUGE-L and F-BERT on the Aya Evaluation Suite, but reduces accuracy on European-IFEval by 29.8% relative to the unadapted baseline. In contrast, adaptation with EuroAlpaca improves accuracy by 12.9% over the same baseline, reversing the degradation caused by direct MT, while also achieving the highest ROUGE-L and F-BERT scores on Aya. These results show that preserving task semantics is essential for multilingual instruction tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑