arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

手语手形的零样本跨语言识别

Zero-Shot Cross-Lingual Recognition of Sign Language Handshapes

Marcel Granero-Moya, Carolina del Corral Farrarós, Gloria Haro, Coloma Ballester, Ricardo Marques

arXiv 2609.18772首次发表:更新:

发表机构

Universitat Pompeu Fabra(庞培法布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个零样本跨语言手形识别框架,利用音系特征分解从ASL迁移至LSC,在37手形基准上达到80%特征准确率,为低资源手语提供无标签迁移方案。

AI 中文摘要

手语处理技术在美国手语(ASL)等高资源语言中发展迅速,但世界上大多数手语缺乏新方法所需的音系标注。我们提出了首个用于手形识别的零样本跨语言框架,实现从美国手语到加泰罗尼亚手语(LSC)的迁移。我们的方法利用手形分解为五个音系特征——所选手指、弯曲度、展开度、拇指位置和拇指接触——这些特征在两种语言中共享,通过复合音系距离度量从预测特征中解码LSC手形。我们评估了三种架构(MLP、SL-GCN、SHuBERT),这些架构在两个ASL语料库(PopSign、Sem-Lex)上训练,并针对一个包含37种手形、单一手语者的LSC基准进行测试。一旦录制格式差异得到协调,零样本迁移被证明是可行的,达到了80.0%的音系特征准确率和54.5%的预期手形准确率。因此,音系分解为将手语技术扩展到低资源语言提供了一座桥梁,而无需任何目标语言的视频训练标签。

英文摘要

Sign language processing advances rapidly for high-resource languages such as American Sign Language (ASL), yet most of the world's sign languages lack the phonological annotations new methods require. We present the first zero-shot cross-lingual framework for handshape recognition, transferring from ASL to Catalan Sign Language (LSC). Our approach leverages the decomposition of handshapes into five phonological features -- selected fingers, flexion, spread, thumb position, and thumb contact -- shared across both languages, to decode LSC handshapes from predicted features via a composite phonological distance metric. We evaluate three architectures (MLP, SL-GCN, SHuBERT) trained on two ASL corpora (PopSign, Sem-Lex) against a 37-handshape, single-signer LSC benchmark. Zero-shot transfer proves viable once recording-format disparities are harmonized, reaching 80.0% phonological feature accuracy and 54.5% expected handshape accuracy. Phonological decomposition thus offers a bridge for extending sign language technologies to low-resource languages without any target-language video training labels.

CommentsAccepted at the Workshop on Sign Language Processing (WSLP), EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑