发表机构
ADAPT Centre, Munster Technological University; Nimbus Research Centre, Munster Technological University; EEECS, Queen’s University Belfast(ADAPT中心,芒斯特理工大学; 宁巴斯研究中心,芒斯特理工大学; 电子、电气与计算机工程学院,贝尔法斯特女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过手语者折交叉验证揭示手语者依赖评估高估手语翻译模型性能,提出采用手语者无关协议、句子不重叠划分及双重报告以提升评估可靠性。
AI 中文摘要
手语翻译已随着深度学习取得进展,然而评估在很大程度上仍是手语者依赖的,训练/开发/测试集中存在重叠的手语者。这引发了关于模型是真正泛化还是依赖手语者特定规律的问题。我们在GFSLT-VLP、GASLT和SignCL这三个领先的、公开可用的无词汇注释手语翻译模型上,使用CSL-Daily和PHOENIX14T数据集进行了手语者折交叉验证。在手语者无关评估下,性能急剧下降:在PHOENIX14T上,GFSLT-VLP的BLEU-4从21.44降至3.59,ROUGE-L从42.49降至11.89;GASLT从15.74降至8.26;SignCL从22.74降至3.66。我们还观察到,在CSL-Daily中许多目标句子由多个手语者执行,因此常见的划分可能将相同句子同时置于训练集和测试集中,通过奖励对重复句子的回忆而非真正的泛化来虚增绝对分数。这些发现表明,手语者依赖的评估可能大幅高估手语翻译能力。我们建议:(1)采用手语者无关的评估协议以确保对未见手语者的泛化;(2)重构数据集,包含明确的手语者无关、句子不重叠的划分,以进行一致的基准测试;(3)同时报告手语者依赖和手语者无关的结果,并附上训练-测试句子重叠情况,以提高透明度和可比性。
英文摘要
Sign Language Translation has advanced with deep learning, yet evaluations remain largely signer-dependent, with overlapping signers across train/dev/test. This raises concerns about whether models truly generalise or instead rely on signer-specific regularities. We conduct signer-fold cross-validation on GFSLT-VLP, GASLT, and SignCL, three leading, publicly available, gloss-free SLT models, on CSL-Daily and PHOENIX14T. Under signer-independent evaluation, performance drops sharply: on PHOENIX14T, GFSLT-VLP falls from BLEU-4 21.44 to 3.59 and ROUGE-L 42.49 to 11.89; GASLT from 15.74 to 8.26; and SignCL from 22.74 to 3.66. We also observe that in CSL-Daily many target sentences are performed by multiple signers, so common splits can place identical sentences in both training and test, inflating absolute scores by rewarding recall of recurring sentences rather than genuine generalisation. These findings indicate that signer-dependent evaluation can substantially overestimate SLT capability. We recommend: (1) adopting signer-independent protocols to ensure generalisation to unseen signers; (2) restructuring datasets to include explicit signer-independent, sentence-disjoint splits for consistent benchmarking; and (3) reporting both signer-dependent and signer-independent results together with train-test sentence overlap to improve transparency and comparability.
Comments13 pages. Published in Findings of the Association for Computational Linguistics: EMNLP 2025
Journal refFindings of EMNLP 2025, pages 18379-18391, 2025
DOI:10.18653/v1/2025.findings-emnlp.997