arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FACET 在 WMT 2026 自动翻译质量评估任务中的应用

FACET at WMT 2026 Automated Translation Quality Evaluation Task

Ahrii Kim, Chanjun Park, Seong-heum Kim

arXiv 2610.00096首次发表:更新:

发表机构

AI-Bio Convergence Research Inst.; Soongsil University(AI-生物融合研究所; 崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FACET 是一种无参考的翻译质量评估方法,通过流畅性、准确性和一致性三个环节分解评估,仅使用固定模型提示三次,无需训练组件,在 WMT26 任务中排名后编辑人工翻译第一,一致性环节对分数影响小。

AI 中文摘要

机器翻译中的不同错误类型需要不同的证据。意义是否得到保留只能根据源文本来判断,而目标文本是否结构良好,或者是否一致地命名某个实体,则可以仅根据目标文本来判断。我们提出了 FACET,这是我们提交给 WMT26 自动翻译质量评估任务的无参考系统,它将评估分解为流畅性、准确性和一致性三个环节,并且每个环节只获得其错误类型所需的上下文。一个固定的模型被提示三次,合并后的错误跨度产生三个任务输出:错误跨度、质量分数和无错误标签,且不包含任何训练组件。我们还提交了 FACET-C,它省略了一致性环节。在没有黄金标签的情况下,我们描述了 FACET 的预测特征。其系统排名将后编辑的人工翻译排在首位,而一致性环节改变了约十分之一的片段分数,同时几乎不改变排名。

英文摘要

Different error types in machine translation require different evidence. Whether meaning is preserved can be judged only against the source, while whether the target is well-formed, or whether it names one entity consistently, can be judged from the target alone. We present FACET, our reference-free submission to the WMT26 Automated Translation Quality Evaluation Task, which decomposes evaluation into Fluency, Accuracy, and Consistency passes and gives each pass only the context its error type requires. A single fixed model is prompted three times, and the merged error spans yield the three task outputs, error spans, quality scores, and error-free labels, with no trained components. We also submit FACET-C, which omits the Consistency pass. Without gold labels, we characterize the predictions of FACET. Its system rankings place post-edited human translation first, and the Consistency pass changes about a tenth of segment scores while leaving the ranking nearly unchanged.

CommentsAccepted at WMT 2026 (shared task system paper)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑