arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究时间运动特征用于姿态到文本的印度手语翻译

Investigating Temporal Motion Features for Pose-to-Text Indian Sign Language Translation

Manav Dhamecha, Praveen Kumar Chandaliya, Pruthwik Mishra

arXiv 2609.12993首次发表:更新:

发表机构

Sardar Vallabhbhai National Institute of Technology(萨达尔·瓦拉巴伊国家技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对WSLP 2026共享任务,探究T5模型规模与显式运动特征对姿态到文本印度手语翻译的影响,发现T5-small加运动特征显著提升BLEU,系统排名第5。

AI 中文摘要

我们研究了预训练T5模型规模和显式运动特征对WSLP 2026共享任务中姿态到文本印度手语翻译(SLT)的影响。姿态序列通过一个轻量级姿态编码器投影到T5的嵌入空间,并对完整模型进行微调以生成英文文本。本工作使用的共享任务数据包括包含5,334个样本的测试集和包含5,257个样本的验证集。我们比较了T5-small、T5-base和T5-large,并额外引入了一种运动增强变体T5-small + Motion,该变体在输入表示中添加了显式的帧间姿态差异。在仅空间模型中,T5-small取得了最佳的BLEU和ROUGE分数,而T5-large获得了最高的chrF分数。为T5-small添加运动特征带来了我们研究中观察到的最大单次改进,显著提高了相对于仅空间基线的BLEU分数,使其在该指标上成为整体最强的模型。我们提交的系统在官方WSLP 2026 SLT测试排行榜上排名第5。源代码和训练好的模型已在GitHub和HuggingFace上公开。

英文摘要

We investigate the effect of pretrained T5 model scale and explicit motion features on pose-to-text Indian Sign Language Translation (SLT) for the WSLP 2026 Shared Task. Pose sequences are projected into the embedding space of T5 through a lightweight pose encoder, with the complete model fine-tuned to generate English text. The shared task data used for this work consists of a test set with 5,334 examples and a validation set with 5,257 examples. We compare T5-small, T5-base, and T5-large, and additionally introduce a motion-augmented variant, T5-small + Motion, that adds explicit frame-to-frame pose differences to the input representation. T5-small achieves the best BLEU and ROUGE scores among the spatial-only models, while T5-large obtains the highest chrF score. Augmenting T5-small with motion features yields the largest single improvement observed in our study, substantially improving BLEU over the spatial-only baseline and making it the strongest model overall on this metric. Our submitted system ranked 5th on the official WSLP 2026 SLT testing leaderboard. The source code and trained models are publicly available on GitHub and HuggingFace.

Comments5 pages, 1 table and 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑