发表机构
Variable Energy Cyclotron Centre; Homi Bhabha National Institute; Ramakrishna Mission Vivekananda Educational and Research Institute(可变能量回旋加速器中心; 霍米·巴巴国家研究所; 罗摩克里希纳使命维韦卡南达教育与研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了基于HamNoSys的160类14.4万张RGB图像的手语手形数据集,评估了ResNet-18等四类模型,建立了可复现基准,为手语技术开发提供支持。
AI 中文摘要
研究背景与问题:细粒度手形识别支持计算手语转录、识别与翻译,但面向语音学定义的、考虑手语使用者差异的视觉 inventory(此处保留术语)数据集仍有限。方法:本研究提出基于独立于语言的汉堡符号系统(HamNoSys)的基准资源。从15名参与者处收集了144000张RGB图像,构成平衡数据集,对应HamNoSys 4官方手形图表定义的160类手形。评估了两种外观模型:ResNet-18和ViT-B/16;同时评估了两种基于手部关键点的模型:图卷积网络和XGBoost。采用了按类别分层的受试者依赖划分,以及15折留一受试者(LOSO)协议,还在LSWH100和ASL手指拼写数据集A上评估了相同模型家族以进行外部验证。结果:受试者依赖基准在所有四个模型家族上建立了可复现的参考性能,而LOSO评估显示当需要对未见过的参与者进行泛化时,识别性能大幅下降;在ASL手指拼写数据集A上,平均LOSO Top-1准确率为82.20%至87.40%。结论:所记录的数据采集、整理及互补评估协议,为细粒度孤立手形研究及开发更易用的手语技术提供了可复现的资源。
英文摘要
Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg Notation System (HamNoSys). Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 were evaluated as appearance-based models, while a graph convolutional network and XGBoost were evaluated from hand landmarks. Both a class-stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent benchmarks established reproducible reference performance across all four model families, whereas LOSO evaluation exposed a substantial reduction when recognition was required to generalise to unseen participants. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% to 87.40%. Conclusion: The documented acquisition, curation, and complementary evaluation protocols pro-vide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies.