发表机构
Kwansei Gakuin University; Georgia Institute of Technology; Keio University; Technical University of Clausthal; The University of Tokyo(关西学院大学; 佐治亚理工学院; 庆应义塾大学; 克劳斯塔尔工业大学; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对日语手语缺乏大规模多手势者数据集的问题,提出JSL-DC数据集,其含36700个视频,受其描述启发的模型在易混淆子集上性能超现有方法9.8%,将加速手语识别研究。
AI 中文摘要
有效的手语(SL)习得对失聪儿童至关重要,但95%的失聪儿童出生于通常不熟练掌握手语的听力正常父母家庭。手语识别技术可开发学习工具,帮助父母与孩子沟通。然而,日语手语(JSL)缺乏大规模多手势者数据集,阻碍了能泛化到新用户的模型开发。为解决这一缺口,我们推出JSL-DC,它是视频数量最多的JSL数据集,包含19名手势者的36700个视频。整个过程以失聪者为中心:由失聪者和科达(Coda)语言学家选定包含270个JSL词汇的词表,以促进亲子沟通;所有参与者为日常使用JSL的失聪者;数据经过失聪语言学家的两阶段审核流程。此外,我们提供用于区分易混淆手势的语言学家推导描述。我们证明,受该描述启发的所提模型在易混淆子集上的表现比最先进的识别方法高出9.8%。该数据集及启发新模型的语言描述将以CC-BY 4.0许可发布,以加速手语识别领域的研究。
英文摘要
Effective sign language (SL) acquisition is crucial for deaf children, yet 95% are born to hearing parents who often lack proficiency in SL. SL recognition can power learning tools to help parents communicate with their children. However, Japanese Sign Language (JSL) lacks large-scale, multi-signer datasets, hindering the development of models that can generalize to new users. To address this gap, we introduce JSL-DC, the largest JSL dataset by video count, comprising 36.7K videos from 19 signers. The entire process was Deaf-centric: the lexicon comprising 270 JSL words was selected by Deaf and Coda linguists to facilitate parent-child communication, all participants were Deaf individuals who use JSL daily, and the data underwent a two-stage review process involving Deaf linguists. Moreover, we provide linguist-derived descriptions for distinguishing confusable signs. We demonstrate that the proposed model inspired by the descriptions outperforms state-of-the-art recognition methods by 9.8% on the confusable subset. The dataset, along with its linguistic description that inspires new models, will be released under a CC-BY 4.0 license to accelerate research in SL recognition.