发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以古文汉英翻译为测试案例,引入诊断框架探查现有翻译评估指标的可靠性,发现MetricX24整体表现最佳,凸显需开发更适配历史文化独特场景的翻译评估指标。
AI 中文摘要
尽管大语言模型能出色地翻译部分历史语言,但数字人文工作流中因缺乏可靠评估,其实用性受限。本研究以古文汉英翻译为测试案例,探究现有针对现代语言开发的自动评估指标在此场景下是否可靠。我们引入基于最小对的诊断框架,该框架捕捉学术使用中显著的错误类型,探查参考依赖型与参考自由型指标的错误敏感性及对有效变异的容忍度。研究发现所有指标均存在盲区,不过MetricX24整体表现最佳,结果凸显针对历史文化独特的翻译场景,亟需更稳健、可解释的指标。
英文摘要
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.