arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言大语言模型生成文本中的翻译腔研究

An Investigation of Translationese in the Generations of Multilingual Large Language Models

Maria Valentini, Téa Wright, Julisa Granados, Eliana Colunga, Katharina von der Wense

arXiv 2608.17399首次发表:更新:

发表机构

University of Colorado Boulder; University of California Berkeley; Johannes Gutenberg University Mainz(科罗拉多大学博尔德分校; 加州大学伯克利分校; 美因茨约翰内斯·古腾堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文探究多语言大语言模型生成文本是否存在翻译腔,对比直接翻译的翻译腔差异,通过多种方法评估MLLM生成文本的翻译腔特征。

AI 中文摘要

从其他语言翻译而来的文本往往带有翻译的痕迹,因此常被称为“翻译腔(translationese)”。多语言大语言模型(MLLMs)可生成多种语言的文本,但目前仍不清楚其生成文本是否类似内部翻译(来自英语或其他语言),进而产生翻译腔。本文提出两个研究问题:(1)MLLMs生成的文本是否类似翻译腔?(2)MLLMs产生的翻译腔与直接翻译产生的翻译腔有何差异?本文利用已有的翻译文本指标,评估最先进的MLLMs在五种语言中生成的文本,同时与非翻译文本和人工编写的基准进行比较,以分离翻译腔与其他干扰因素。通过使用高精度分类模型、对单个语言特征的方差分析,以及对德语和西班牙语两种语言子集的人工标注收集,本文评估了MLLM生成文本的翻译腔含量,并研究了区分MLLM生成文本与典型翻译相关干扰的关键特征。

英文摘要

Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs' generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.

CommentsAccepted to COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑