发表机构
University of Trento; Fondazione Bruno Kessler(特伦托大学; 布鲁诺·凯斯勒基金会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出以SMPL-X运动为核心、手语词汇引导局部对齐和视觉蒸馏的手语检索框架,在CSL-Daily和PHOENIX-2014T上实现领先的双向检索性能。
AI 中文摘要
手语与文本的对齐仍然是文本驱动手语理解中的一个基本挑战。现有方法主要依赖外观为主的RGB表示,这会将运动语义与视觉变化纠缠在一起,导致运动与语言的对齐产生歧义。在本文中,我们在一个结构化的运动学空间中重新构建手语与文本的对齐,并提出一个以运动学为中心的框架,采用3D SMPL-X运动作为主要表示。通过在一个统一的运动空间中显式建模手语的运动动力学,我们的方法减少了对表观信号的依赖,并产生语义上更一致的表示。为了捕捉手语的组合特性,我们引入了一种手语词汇引导的局部对齐机制,利用手语词汇的时间跨度作为弱监督,将连续运动分解为连贯的片段,并建立细粒度的运动-文本对应关系,从而减少在连续手语中定位词级语义的歧义。此外,我们开发了一种视觉蒸馏策略,其中RGB信号在训练期间作为特权监督提供互补的上下文线索,而在推理时完全移除。在标准基准上的大量实验表明,我们的方法在CSL-Daily上实现了最先进的双向检索性能,并在PHOENIX-2014T上取得了有竞争力的结果。这些结果突显了运动学表示和显式局部基础对于手语-文本对齐的有效性。
英文摘要
Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sign language-text alignment in a structured kinematic space and propose a kinematics-centric framework that adopts 3D SMPL-X motion as the primary representation. By explicitly modeling the kinematic dynamics of signing in a unified motion space, our approach reduces reliance on appearance signals and yields more semantically consistent representations. To capture the compositional nature of sign language, we introduce a gloss-guided local alignment mechanism that leverages gloss temporal spans as weak supervision to decompose continuous motion into coherent segments and establish fine-grained motion-text correspondences, thereby reducing ambiguity in localizing word-level semantics in continuous signing. Furthermore, we develop a visual distillation strategy, where RGB signals serve as privileged supervision during training to provide complementary contextual cues, while being completely removed at inference time. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art bidirectional retrieval performance on CSL-Daily and competitive results on PHOENIX-2014T. These results highlight the effectiveness of kinematic representations and explicit local grounding for sign language-text alignment.
CommentsAccepted by ACM Multimedia 2026 (ACM MM 2026)