发表机构
School of Computer Science, Shanghai Jiao Tong University; School of Foreign Languages, East China University of Science and Technology; Mejiro University(上海交通大学计算机科学与工程系; 华东理工大学外国语学院; 目白大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究以《红楼梦》构建数据集,系统探究基于大语言模型的机器翻译系统中文化负载翻译的挑战,揭示任务、人工评估及自动评估方面的问题,为文化导向的翻译研究提供见解。
AI 中文摘要
文化负载翻译给机器翻译带来独特挑战,因为意义深深嵌入社会文化语境而非表面语言形式。虽大语言模型使机器翻译系统在许多场景达到类人质量,但其处理文化负载表达的能力仍未充分探索。本研究系统调查基于大语言模型的机器翻译系统中文化负载翻译带来的挑战。我们从具有文化代表性的《红楼梦》语料库构建中日双语数据集,含500个不同文化类别的片段。通过综合评估协议,揭示三个主要挑战:前沿大语言模型在文化负载内容上表现有差距;评估者背景导致翻译判断存在重大分歧;广泛使用的指标无法可靠评估此任务的翻译质量。这些发现可为计算科学和语言学中面向文化的翻译研究提供有价值的见解。
英文摘要
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally loaded translation in LLM-based MT systems. We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of the Red Chamber, containing 500 segments across diverse cultural categories. Using a comprehensive evaluation protocol, we reveal three main challenges: (1) task challenges, where frontier LLMs exhibit notable performance gaps and struggle with culturally loaded content; (2) human evaluation challenges, where evaluator backgrounds lead to substantial disagreement in translation judgments; and (3) automatic evaluation challenges, where widely used metrics fail to reliably assess translation quality for this task. These findings may offer valuable insights for culture-oriented translation research in both computational science and linguistics.