arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TransClean:一个用于检测和提取大语言模型输出中干净翻译的基准

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

Shenbin Qian, Yves Scherrer

arXiv 2609.11399首次发表:更新:

发表机构

University of Oslo(奥斯陆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM翻译输出中的噪声问题,构建了包含9900对噪声/干净翻译的基准TransClean,并评估了两种提取方法,为评估和改进翻译清洁度提供了首个系统框架。

AI 中文摘要

大语言模型(LLMs)越来越多地用于机器翻译,然而其输出往往包含除翻译本身之外的额外文本,例如语言标签、解释或双语重复,我们称之为翻译噪声。尽管这一问题普遍存在,但缺乏专门的基准和系统性研究。我们分析了来自12个LLM在22个语言对(LPs)上的超过790,000个翻译输出,识别出12种反复出现的噪声模式,并将其分为格式噪声和内容噪声。基于观察到的模式,我们构建了TransClean,一个包含9,900对噪声和干净翻译输出的受控基准,其中包含8,800个合成生成的实例和1,100个手动整理的权威实例。我们在TransClean基准上评估了两种提取方法:1)一种基于跨度的提取方法,利用翻译质量估计模型进行跨度检测;2)一种基于LLM的提取方法,提示LLM隔离出翻译。我们的基准和分析提供了第一个系统框架,用于评估和改进LLM翻译输出的清洁度。

英文摘要

Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevalence, this problem lacks dedicated benchmarks and systematic study. We analyze over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identify 12 recurring noise patterns, which we group into formatting and content noise. Building on the observed patterns, we construct TransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs, comprising 8,800 synthetically generated instances and 1,100 manually curated authentic instances. We evaluate two extraction approaches on the TransClean benchmark: 1) a span-based extraction method leveraging translation quality estimation models for span detection, and 2) an LLM-based extraction method that prompts an LLM to isolate the translation. Our benchmark and analysis provide the first systematic framework to evaluate and improve the cleanliness of LLM translation outputs.

CommentsAccepted to the Eleventh Conference on Machine Translation (WMT26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑