arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合语义与重建之间的差距:统一手语翻译与生成

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie

arXiv 2608.09045首次发表:更新:

发表机构

State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出统一框架Uni-SLTP,解决手语翻译与生成的模态映射反向难题,在公共数据集上实现手语生成更优运动精度且保持翻译竞争力。

AI 中文摘要

近期手语(SL)研究的进展显示出将多个手语理解(SLU)子任务(如孤立手语识别(ISLR)、连续手语识别(CSLR)和手语翻译(SLT))统一在单个框架内的趋势,并取得了显著进展。与此同时,从文本生成手语序列的手语生成(SLP)也受到越来越多的关注。这自然提出了一个重要问题:能否将手语理解和生成为单个框架?与统一SLU子任务相比,这个问题的挑战性要大得多。现有的SLU任务在很大程度上共享相同的映射方向,即从手语输入到语言输出,而SLT和SLP则处于手语-文本映射的相反方向。因此,统一框架必须解决两个关键挑战:(1)通过支持语言抽象和运动重建的共享手语分词器,弥合连续手语运动与离散文本标记之间的模态差距;(2)学习单个条件自回归模型,该模型可以以手语或文本作为输入,并生成对应相反模态的目标序列。为此,我们提出了Uni-SLTP,这是一个用于SLT和SLP的统一框架,包含两个关键组件:(1)共享手语分词器,将手语序列转换为离散标记和潜在表示,同时捕捉语义和重建信息;(2)统一自回归生成模型,将两个任务均表述为条件序列生成。在广泛使用的公共数据集上的实验表明,Uni-SLTP在SLP中实现了更优的运动准确性,同时保持了具有竞争力的SLT性能。

英文摘要

Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress. Meanwhile, sign language production (SLP), which generates sign sequences from text, has also attracted growing attention. This naturally raises an important question: can sign language understanding and production be unified within a single framework? Compared with unifying SLU subtasks, this problem is substantially more challenging. Existing SLU tasks largely share the same direction of mapping, namely from sign inputs to linguistic outputs, whereas SLT and SLP lie in opposite directions of sign-text mapping. A unified framework must therefore address two key challenges: (1) bridging the modality gap between continuous sign motions and discrete text tokens through a shared sign tokenizer that supports both linguistic abstraction and motion reconstruction; and (2) learning a single conditional autoregressive model that can take either sign or text as input and generate the corresponding target sequence in the opposite modality. To this end, we propose Uni-SLTP, a unified framework for SLT and SLP with two key components: (1) a shared sign tokenizer that converts sign sequences into discrete tokens and latent representations, capturing both semantic and reconstructive information; and (2) a unified autoregressive generation model that formulates both tasks as conditional sequence generation. Experiments on widely used public datasets show that Uni-SLTP achieves superior motion accuracy for SLP while maintaining competitive SLT performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑