arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16344cs.CL

IndicQE-APE:印度语言质量估计与自动后编辑基准

IndicQE-APE: A Consolidated Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Orăsan, Chrysoula Zerva, Ric… 展开作者

Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Orăsan, Chrysoula Zerva, Ricardo Rei, Frédéric Blain, André F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建了涵盖9个方向对的IndicQE-APE基准,对LLM、COMET指标及APE系统进行QE与APE基准测试,揭示质量信号冲突等因素对文本段排名的影响。

中文摘要 AI 辅助

印度语言的质量估计(QE)和自动后编辑(APE)数据分散在不同的发布版本中,因此没有单一资源能同时支持跨任务和语言对的训练与评估。我们将WMT 2020至2024年的共享任务系列与扩展的英语-马拉雅拉姆语资源整合为\textit{IndicQE}:共包含9个方向对的126754个实例,同一文本段上对齐了多达4种标签类型,分别为直接评估结果、人工后编辑内容、词级OK/BAD标签以及错误说明,测试集按4个难度轴分层构建。基于该基准,我们对6个采用提示学习的大型语言模型(LLM)和3个COMET指标进行了段级QE基准测试,对3个系统进行了APE基准测试。其中2个难度轴部分基于直接评估定义,并从中选取了压缩子集,因此每个轴与来自同一语言对、具有相同分数分布的对照组进行比较。只有1个轴通过了对照组检验:对于所有9个系统和承载该轴的7个语言对,整体与词级质量信号冲突的文本段,其排名低于同一语言中得分相同的文本段。未使用对照组时看似难度第二高的标注者分歧,在使用对照组后无影响。少样本提示使每个模型的相关性和输出格式合规性均损失不超过34亿(≤3.4B)。语言内准确性无法使不同语言对的分数具有可比性:在3个训练指标中,语言内相关性最佳的指标在合并语言对时损失最大。该基准及代码将发布。

英文摘要

Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020-2024 shared-task lineage with an extended English-Malayalam resource into IndicQE-APE: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level tags and an error explanation, and a test set stratified over four difficulty axes. We benchmark six prompted LLMs and three COMET metrics on segment-level QE, and three systems on APE. Two of the axes are defined partly on direct assessment and select a compressed slice of it. Segments whose segment-level and token-level signals disagree are ranked below equally scored segments of the same language. Four-shot prompting costs every model at or below $3.4$B both correlation and output-format compliance. Unedited MT beats every APE system we run on three of the four pairs. The benchmark (https://huggingface.co/datasets/surrey-nlp/IndicQE-APE) and code (https://github.com/surrey-nlp/IndicQE-APE) are released.

发表机构

  • IIT Bombay(印度理工学院孟买分校)
  • University of Oslo(奥斯陆大学)
  • Lancaster University(兰卡斯特大学)
  • INESC-ID
  • Instituto de Telecomunicações(电信研究所)
  • Instituto Superior Técnico(高等技术学院)
  • University of Lisbon(里斯本大学)
  • Sword Health
  • Tilburg University(蒂尔堡大学)
  • Zoom Communications(Zoom通信公司)
  • Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)
  • Apple(苹果公司)
  • Bodhan AI
  • IIT Madras(印度理工学院马德拉斯分校)
  • University of Surrey(萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑