发表机构
Home Lab(家庭实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文以9种欧盟语言为对象,对比3款本地LLM与专用MT基线,研究提示范围、示例选择对机器翻译的影响,发现专用MT系统仍最优,嵌入检索对强LLM效果较好,语族范围提示对强LLM可行但存在小型模型缺陷。
AI 中文摘要
大语言模型(LLM)正日益被用作通用翻译系统,但其评估通常仅采用单一提示形式:将单个源句子翻译成单个目标语言。实际应用中,用户可能要求翻译成单个目标语言、多种相关语言,或基于示例的翻译。本文将提示范围与示例选择作为本地LLM机器翻译的实验变量展开研究。针对9种欧盟官方语言,在完整的FLORES devtest拆分数据集上,评估英语到罗曼语族、英语到日耳曼语族的翻译性能。对比3款本地指令微调LLM:llama3.2:3b、mistral:latest、qwen2.5:14b,以及来自OPUS-MT和NLLB-200的专用机器翻译(MT)基线。测试零样本提示与k=5的少样本提示,示例选择方式包括随机选择、词汇相似性选择、嵌入相似性选择;同时对比单目标提示与JSON格式的语族范围提示(要求一次性翻译语族内所有语言)。结果显示,专用MT系统整体仍表现最优,尤其在日耳曼语族翻译中;少样本提示对mistral:latest和qwen2.5:14b有帮助,但对llama3.2:3b有负面影响;嵌入检索对性能较强的LLM平均效果最佳,但其相比随机和词汇示例的优势有限;语族范围提示对性能较强的本地LLM可行,但会暴露出小型模型在结构化输出上的缺陷。这些发现表明,评估LLM翻译不仅需关注语言对与指标,还应考虑提示范围、检索策略及多目标合规性。
英文摘要
Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt shape: translate one source sentence into one target language. In practice, users may ask for one target language, for several related languages at once, or for translations conditioned on examples. This paper studies prompt scope and demonstration selection as experimental variables for local LLM machine translation. We evaluate English-to-Romance and English-to-Germanic translation on the full FLORES devtest split for nine official European Union languages. We compare three local instruction-tuned LLMs, llama3.2:3b, mistral:latest, and qwen2.5:14b, against dedicated MT baselines from OPUS-MT and NLLB-200. We test zero-shot prompting and k=5 few-shot prompting with random, lexical-similarity, and embedding-similarity demonstration selection. We also compare single-target prompts with JSON-formatted family-scope prompts that request all languages in a family at once. Results show that dedicated MT systems remain strongest overall, especially for Germanic languages. Few-shot prompting helps mistral:latest and qwen2.5:14b, but hurts llama3.2:3b; embedding retrieval is best on average for the stronger LLMs, but its advantage over random and lexical examples is modest. Family-scope prompting is feasible for stronger local LLMs but exposes structured-output failures in smaller models. These findings motivate evaluating LLM translation not only by language pair and metric, but also by prompt scope, retrieval strategy, and multi-target compliance.