我们能否在没有人工参考的情况下对古典文本的LLM翻译错误进行分类?源文本新颖性、GEMBA评分与通过巴利语到英语翻译的预算审查
Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation
浏览论文内容
中文总结 AI 辅助
本研究通过巴利语到英语翻译测试无参考错误分类,发现无参考GEMBA评分是最强信号,审查前10%可捕获81.6%错误,并提出结合多种信号的预算审查工作流程。
中文摘要 AI 辅助
随着大型语言模型成为古典文本的熟练翻译者,一个关键挑战是在没有人工参考的情况下决定哪些输出需要专家审查。本研究通过巴利语到英语的翻译测试了无参考错误分类。三个LLM翻译了15,493个段落。比较了五种信号:源文本新颖性、源文本-候选文本嵌入距离、同行翻译分歧、英语到巴利语回译以及无参考GEMBA评分。信号在3,000项参考信息LLM裁决样本上进行了校准,并与500项作者裁决锚点进行了核对。人工参考仅用于校准和验证;从未用于计算风险信号。源文本新颖性是有用的源端风险先验,但不是针对每个候选的错误检测器。同行分歧和回译提供了次要信号。最强的方法是由通常被认为比翻译者更强的模型小组进行无参考GEMBA评分:按GEMBA风险审查前10%在校准集中捕获了81.6%的小组多数错误。GEMBA在作者锚点上仍然是最佳的无参考信号。同一级别的小组,排除自我评分,仍然有用但表现较差,表明评估者的实力不仅仅取决于提示。提出了一种预算工作流程,结合源文本新颖性、同行分歧和更强的候选感知判断来分配人工审查。向其他古典语言(包括拉丁语、古希腊语和梵语)的迁移仍有待测试。
英文摘要
As large language models become capable translators of classical texts, a key challenge is deciding which outputs need expert review when no human reference exists. This study tests reference-free error triage through Pali-to-English translation. Three LLMs translated 15,493 passages. Five signals were compared: source novelty, source-candidate embedding distance, peer-translation disagreement, English-to-Pali backtranslation, and no-reference GEMBA scoring. Signals were calibrated on a 3,000-item reference-informed LLM-adjudicated sample and checked against a 500-item author-adjudicated anchor. Human references supported calibration and validation only; they were never used to compute the risk signals. Source novelty was a useful source-side risk prior but not a per-candidate error detector. Peer disagreement and backtranslation provided secondary signal. The strongest method was no-reference GEMBA scoring by a panel of models generally regarded as stronger than the translators: reviewing the top 10% by GEMBA risk captured 81.6% of panel-major errors in the calibration set. GEMBA also remained the best reference-free signal against the author anchor. A same-tier panel, with self-scoring excluded, remained useful but performed worse, indicating that evaluator strength matters beyond the prompt alone. A budgeted workflow is proposed, combining source novelty, peer disagreement, and stronger candidate-aware judging to allocate human review. Transfer to other classical languages, including Latin, Ancient Greek, and Sanskrit, remains to be tested.