arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03788cs.CVcs.CL

反向手语词典:通过视频字幕生成与描述检索实现连续手语的开放词汇识别

A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval

  • The University of Tokyo(东京大学)
  • National Institute of Informatics(情报学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Santiago Poveda-Gutiérrez, Hideki Nakayama, Mayumi Bono

中文总结 AI 辅助

该研究提出无需 gloss 监督的反向手语词典,通过开放权重视觉语言模型生成手语描述、多语言句子编码器检索,实现日本手语的开放词汇识别,可见与未见类检索均获显著提升。

中文摘要 AI 辅助

孤立手语识别(ISLR)传统上被视为对 gloss 标签的闭集分类,无法泛化到训练中未出现的手语,且每次部署都依赖带有 gloss 注释的词汇表。相反,我们通过两种方法识别从连续手语中提取的手语:(1)用开放权重的视觉语言模型将手语级片段生成为关于发音的自由形式过程描述;(2)用多语言句子编码器从目标描述词汇表中检索最接近的条目:即一种反向手语词典,无需 gloss 监督且支持开放词汇。在带有过程描述注释的日本手语(JSL)对话语料库的 1300 个手语级片段上(目标词汇表有 503 个条目,top-10 随机概率下限为 2%),对字幕生成器进行微调显著提升了可见类别的检索:语言和视觉塔微调将可见类别的 top-10 检索从 4.5%(未训练)提升至 49%,在闭集分类器(I3D)可评估的三个测试集中的两个上,其性能与标准监督闭集分类器无统计学差异。更重要的是,未见类别的检索也比未训练的 pipeline 显著提升(top-10 从 11.5% 升至 21.0%,p=0.0094),而闭集分类器无法参与该场景。匹配器端的经验上限分析显示,多语言句子编码器已能恢复近 100% 的转述黄金描述,表明字幕生成质量存在差距,这是我们未来工作的改进方向。据我们所知,这是首个无需 gloss 监督、基于描述的连续手语开放词汇查找方法,也是针对 JSL 的此类研究。

英文摘要

Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning a sign-level clip into a free-form procedural description of the articulation with an open-weight vision-language model, and (2) retrieving the closest entry from a vocabulary of target descriptions with a multilingual sentence encoder: a reverse sign language dictionary that needs no gloss supervision and admits an open vocabulary. On 1,300 sign-level segments from a Japanese Sign Language (JSL) dialogue corpus annotated with procedural descriptions (against a 2% top-10 chance floor over the 503-entry target vocabulary), fine-tuning the captioner substantially improves seen-class retrieval: language and vision tower fine-tuning raises top-10 retrieval on seen classes from 4.5% (untrained) to 49%, becoming statistically indistinguishable from a standard supervised closed-set classifier (I3D) on two of the three test sets where a closed-set classifier can be evaluated at all. More importantly, unseen-class retrieval also improves significantly over the untrained pipeline (11.5% -> 21.0% top-10, p=0.0094), a regime in which the closed-set classifier cannot participate. A matcher-side empirical upper-bound analysis shows the sentence encoder already recovers close to 100% of paraphrased gold descriptions, locating a gap in captioning quality that we aim to address in future work. To our knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.

补充信息

↑