arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2304.03245cs.CL

大型语言模型能有效利用文档级上下文进行文学翻译,但关键错误依然存在

Large language models effectively leverage document-level context for literary translation, but critical errors persist

  • Manning College of Information and Computer Sciences(曼宁信息与计算机科学学院)
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

机构由 AI 辅助整理,请以论文原文为准。

Marzena Karpinska, Mohit Iyyer

更新

AI总结:

通过严格人工评估发现,Gpt-3.5一次性翻译文学段落比逐句翻译质量更高且错误更少,但仍存在关键错误需人工干预,并公开了相关数据集。

AI中文摘要:

大型语言模型(LLMs)在广泛的句子级翻译数据集上与现有最先进技术不相上下。然而,由于这些场景下的评估成本高且困难,它们翻译段落和文档的能力仍未被探索。我们通过严格的人工评估表明,要求 Gpt-3.5 (text-davinci-003) LLM 一次性翻译整个文学段落(例如来自小说),在18个语言多样化的语言对(例如日语、波兰语和英语之间的互译)中,能产生比标准逐句翻译更高质量的译文。我们的评估耗时约350小时进行标注和分析,通过雇佣精通源语言和目标语言的译员来进行,要求他们提供跨度级别的错误标注以及对哪个系统译文更好的偏好判断。我们观察到,篇章级 LLM 翻译器比句子级方法犯的误译、语法错误和风格不一致更少。尽管如此,关键错误依然频发,包括偶尔的内容遗漏,并且仍需要人类译员的干预以确保作者的声音保持不变。我们公开发布我们的数据集和错误标注,以推动未来关于文档级文学翻译评估的研究。

英文摘要:

Large language models (LLMs) are competitive with the state of the art on a wide range of sentence-level translation datasets. However, their ability to translate paragraphs and documents remains unexplored because evaluation in these settings is costly and difficult. We show through a rigorous human evaluation that asking the Gpt-3.5 (text-davinci-003) LLM to translate an entire literary paragraph (e.g., from a novel) at once results in higher-quality translations than standard sentence-by-sentence translation across 18 linguistically-diverse language pairs (e.g., translating into and out of Japanese, Polish, and English). Our evaluation, which took approximately 350 hours of effort for annotation and analysis, is conducted by hiring translators fluent in both the source and target language and asking them to provide both span-level error annotations as well as preference judgments of which system's translations are better. We observe that discourse-level LLM translators commit fewer mistranslations, grammar errors, and stylistic inconsistencies than sentence-level approaches. With that said, critical errors still abound, including occasional content omissions, and a human translator's intervention remains necessary to ensure that the author's voice remains intact. We publicly release our dataset and error annotations to spur future research on evaluation of document-level literary translation.

补充信息

↑