arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03695cs.CV

SignSeek:学习可迁移的手语词典检索表征

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden

首次发表
浏览论文内容

中文总结 AI 辅助

SignSeek 采用显著性引导发音器掩码的对比学习,在多手语数据集预训练后,在跨语料库手语检索等任务上实现了最优性能,且可零样本泛化到新手语并迁移至其他相关任务。

中文摘要 AI 辅助

手语词典是手语学习者的重要资源,但仅通过查询视频自动从词典中检索手语仍是一个具有挑战性的问题,原因在于不同手语使用者之间存在自然的变异性。现有的手语表征学习方法是为闭集识别而构建的,生成的嵌入无法泛化到检索所需的开放集、独立于手语使用者的场景。SignSeek 通过结合显著性引导的发音器掩码的对比学习来弥合这一差距。对比目标对齐不同使用者间同一 gloss(手语 gloss 指手语的书面标注形式)的手语,而我们的发音器显著性引导掩码(Articulator Saliency-Guided Masking, ASGM)定位每个手语的单个最关键发音器。这驱动了两个互补目标:通过单个发音器观察手语的掩码对比对齐(Masked Contrastive Alignment, MAC)损失,以及从周围时空上下文在潜在空间中重建手语的掩码预测(Masked Prediction, MAP)损失。SignSeek 在多种手语的 26.6 万个样本(约 5700 个 gloss)上进行预训练,在无需任何下游微调的情况下,在 ASL-Citizen、WLASL 和 NMFs-CSL 上的跨语料库检索中达到了新的最先进性能。值得注意的是,它对完全未见过的英国手语(British Sign Language, BSL)实现了零样本泛化,优于专门在 BSL 上训练的方法,并且无缝迁移到孤立手语识别和字幕对齐任务,优于先前基于骨架的方法。

英文摘要

Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a sign from a dictionary, given only a query video, remains a challenging problem due to the natural variability between signers. Existing sign representation learning methods are built for closed-set recognition, producing embeddings that do not generalise to the open-set, signer-independent setting that retrieval demands. \textbf{SignSeek} closes this gap by contrastively learning sign representations with saliency-guided articulator masking. A contrastive objective aligns same-gloss signs across signers, while our Articulator Saliency-Guided Masking (ASGM) pinpoints the single most critical articulator per sign. This drives two complementary objectives, a masked contrastive alignment (MAC) loss that sees the sign through a single articulator and a masked prediction (MAP) loss that reconstructs it in latent space from the surrounding spatio-temporal context. Pretrained on 266K samples ($\sim$5,700 glosses) across multiple sign languages, \textbf{SignSeek} sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning. Strikingly, it achieves zero-shot generalisation to an entirely unseen British Sign Language (BSL), surpassing methods explicitly trained on BSL, and transfers seamlessly to isolated sign recognition and subtitle alignment, outperforming prior skeleton-based methods.

发表机构

  • University of Surrey(萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑