arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03734cs.CLcs.AI

超越BLEU:重新定义手语翻译基准的案例

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

  • Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(萨里大学视觉、语音与信号处理中心)
  • Tavus(塔伍斯公司)

机构由 AI 辅助整理,请以论文原文为准。

Oline Ranum, Edward Fish, Simon Hadfield, Richard Bowden

AI总结:

该研究针对手语翻译(SLT)的BLEU-4指标缺陷,提出基于open-weight-LLM问答的替代评估协议,发现gloss监督系统在Phoenix-2014T上比无gloss系统高9.3分,且该协议与人类排名更契合。

AI中文摘要:

BLEU-4是评估手语翻译(SLT)的标准指标,但口语指标可能无法充分反映手语的熟练程度。SLT的多模态、低资源环境使得模型可以利用虚假相关性和口语先验,而非学习更强大的手语表示。在本文中,我们在Phoenix-2014T和CSL-Daily两个数据集上评估了6个SLT模型的时空理解与BLEU-4之间的关系,表明BLEU-4的提升本身并不能作为手语理解更好的证据。本研究引入了一种受语言学习评估启发的替代方案,采用开放权重大语言模型(open-weight-LLM)问答协议来衡量显著内容的保留情况,该协议与人类排名的一致性更高,且比BLEU-4的 paraphrase不变性高6至7倍。将其应用于SLT时,该协议针对内容迁移,对训练-测试重叠的鲁棒性更强,且呈现出该领域的不同图景:在Phoenix-2014T上,5个无 gloss( gloss指手语 gloss,即手语的书面标注形式)系统的表现基本处于噪声范围内,而gloss监督系统的得分高出9.3个百分点,这一差距是BLEU-4无法察觉的。

英文摘要:

BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning stronger sign representations. In this paper, we evaluate the relationship between spatio-temporal understanding and BLEU-4 across six SLT models on Phoenix-2014T and CSL-Daily, showing that gains in BLEU-4 are not on their own evidence of better sign language understanding. This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation. It aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4. Applied to SLT, this protocol targets content transfer, is more robust to train-test overlap, and gives a different picture of the field: the five gloss-free systems are largely within noise of one another on Phoenix-2014T, while the gloss-supervised system stands 9.3 points higher, a gap invisible to BLEU-4.

↑