发表机构
University of Oxford; École des Ponts, IP Paris, Univ Gustave Eiffel, CNRS(牛津大学; 巴黎路桥学院、巴黎理工学院、古斯塔夫·埃菲尔大学、法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SignMatch方法,通过学习原型结构的手语嵌入空间,实现词典手语与连续手语的匹配,在多数据集、任务及手语间泛化能力强,且优于现有方法。
AI 中文摘要
本文的研究目标是将词典手语视频与连续手语视频中的对应手语进行匹配,其中匹配仅由视觉相似性定义,即手形和相对于身体的运动。为实现这一目标,本文从带有手语标注的连续视频中学习原型结构的手语嵌入空间,每个可学习原型对应一个手语类别。随后将孤立的词典视频映射到该手语空间,从而实现词典示例与连续手语实例之间的匹配。该设计支持通过嵌入相似度直接进行词典引导的手语匹配,且仅使用词典示例即可自然扩展到未见过的手语。在ASL-Citizen词典检索、ChaLearn OSLWL词典到连续手语匹配任务,以及使用BOBSL的CSLR2评估进行自动手语标注的实验表明,该方法在数据集、任务和手语语言间具有强泛化能力。在无特定基准监督的情况下,学习到的表示可有效迁移至美国、英国和西班牙手语,在所有三个基准上均优于现有方法。项目页面:this https URL
英文摘要
The objective of this paper is to match dictionary sign videos to corresponding signs in continuous signing videos, where a match is defined by the visual similarity alone - the handshape and motion relative to the body. To achieve this, we learn a prototype-structured sign embedding space from continuous video annotated with signs, where each learnable prototype corresponds to a sign class. Isolated dictionary videos are then mapped into this sign space, enabling the matching between dictionary exemplars and continuous sign instances. This design supports direct dictionary-guided sign matching through embedding similarity and naturally extends to unseen signs using only dictionary exemplars. Experiments on ASL-Citizen dictionary retrieval, ChaLearn OSLWL dictionary-to-continuous sign matching, and using BOBSL's CSLR2 evaluation for automatic sign annotation demonstrate strong generalisation across datasets, tasks and sign languages. Without benchmark-specific supervision, the learned representation transfers effectively across American, British, and Spanish Sign Languages, outperforming prior methods on all three benchmarks. Project page: https://www.robots.ox.ac.uk/~vgg/research/signmatch/
Comments24 pages, 8 figures, Project page: https://www.robots.ox.ac.uk/~vgg/research/signmatch/