arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VTaMo:用于手语翻译的视频-文本对齐模型

VTaMo: Video-Text Alignment Model for Sign Language Translation

Junyi Hu, Zhewen He, Haomian Huang, Aoxiang Yang, Yi Fang

arXiv 2607.09126首次发表:更新:

发表机构

New York University Abu Dhabi; ChatSign Technology(纽约大学阿布扎比分校; ChatSign科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究手语翻译问题,提出VTaMo框架,通过局部、全局对齐及位置对齐对比学习实现多粒度对齐,在多个数据集实验中展现领先性能,各组件贡献互补。

AI 中文摘要

手语翻译(SLT)将连续的手语视频转换为口语化文本。无词汇方法利用预训练的视觉编码器和语言模型,但仅依赖于来自翻译监督的隐式跨模态对齐。我们提出了VTaMo,这是一个在三个层面引入显式多粒度对齐的框架:(1)通过熵正则化最优传输进行局部对齐,使用可学习的空令牌实现细粒度的帧到令牌对应;(2)通过可学习的正交变换进行全局对齐,通过推土机距离校准嵌入空间几何;(3)位置对齐对比学习以获得判别性令牌级表示。在Phoenix-2014T、CSL-Daily、How2Sign和OpenASL上的实验证明了其始终保持的领先性能,消融实验证实了每个组件的互补贡献。代码可从此https URL获取。

英文摘要

Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation supervision alone. We present VTaMo, a framework that introduces explicit multi-granularity alignment at three levels: (1) local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; (2) global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance; and (3) position-aligned contrastive learning for discriminative token-level representations. Experiments on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL demonstrate consistent state-of-the-art performance, with ablations confirming the complementary contributions of each component. Code is available at https://github.com/junyi2005/vtamo.

Comments18 pages, 5 figures, 8 tables. Accepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑