发表机构
New York University Abu Dhabi; ChatSign Technology(纽约大学阿布扎比分校; ChatSign科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究手语翻译问题,提出VTaMo框架,通过局部、全局对齐及位置对齐对比学习实现多粒度对齐,在多个数据集实验中展现领先性能,各组件贡献互补。
AI 中文摘要
手语翻译(SLT)将连续的手语视频转换为口语化文本。无词汇方法利用预训练的视觉编码器和语言模型,但仅依赖于来自翻译监督的隐式跨模态对齐。我们提出了VTaMo,这是一个在三个层面引入显式多粒度对齐的框架:(1)通过熵正则化最优传输进行局部对齐,使用可学习的空令牌实现细粒度的帧到令牌对应;(2)通过可学习的正交变换进行全局对齐,通过推土机距离校准嵌入空间几何;(3)位置对齐对比学习以获得判别性令牌级表示。在Phoenix-2014T、CSL-Daily、How2Sign和OpenASL上的实验证明了其始终保持的领先性能,消融实验证实了每个组件的互补贡献。代码可从此https URL获取。
英文摘要
Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation supervision alone. We present VTaMo, a framework that introduces explicit multi-granularity alignment at three levels: (1) local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; (2) global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance; and (3) position-aligned contrastive learning for discriminative token-level representations. Experiments on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL demonstrate consistent state-of-the-art performance, with ablations confirming the complementary contributions of each component. Code is available at https://github.com/junyi2005/vtamo.
Comments18 pages, 5 figures, 8 tables. Accepted to ECCV 2026