arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向可部署的孟加拉语手语识别:基于专家验证数据与轻量型注意力模型

Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model

Saad Ahmed, Md Khalid Syfullah

arXiv 2608.06252首次发表:更新:

发表机构

Bangladesh Army University of Science and Technology(孟加拉国陆军科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建专家验证的RSBdSL38数据集,提出轻量型注意力卷积网络,实现高准确率孟加拉语手语识别,模型可在普通智能手机部署,相关资源已公开。

AI 中文摘要

孟加拉国的聋人和听障人士主要通过孟加拉语手语(BdSL)进行交流,在个人设备上实现自动BdSL识别可拓宽其获取教育与服务的渠道。现有系统使用的数据集来自受控场景,且未经过专家验证,同时采用的重型预训练主干网络不适合在设备端使用。本文提出RSBdSL38数据集,包含10874张经专家验证的图像,覆盖全部38种BdSL手势,对应孟加拉语字母表的51个字母,由孟加拉国三所特殊学校的真实手语使用者录制完成。本文提出一种基于注意力的轻量型卷积网络,参数规模为298470,由分组瓶颈残差块、通道与空间注意力模块、多尺度深度手部特征块、双池化层及Swish激活函数构建而成。该模型从零开始训练,准确率达96.37%(5个随机种子的均值为95.72%±0.54%),在相同协议下,与9种基于ImageNet预训练的高效架构中表现最佳的模型相比,准确率仅低1.08个百分点,但参数数量减少8.5至68倍,MACs减少1.3至21.7倍。经重新训练后,该模型在6个公开BdSL基准上的准确率为92.95%至98.33%,在合并语料库上的准确率为97.04%,在BdSL-38上的零样本准确率为76.25%。移除任意一个架构阶段会导致准确率下降7.61至89.30个百分点,而训练配方的影响最多为3.17个百分点。结合删除-插入与权重随机化检验的Grad-CAM分析证实,模型预测依据为手语使用者的手部。在36名手语使用者中留出6名作为独立测试集的情况下,模型准确率为85.18%。将模型量化至0.48 MB后,在普通智能手机上运行时,单张图像推理耗时3.98 ms,内存占用为15.5 MB。综上,RSBdSL38数据集与本文提出的从零开始训练的模型,以远低于预训练主干网络的成本,将基准准确率转化为可部署的易用性;相关数据集、代码与模型已公开。

英文摘要

Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting datasets without expert verification and heavyweight pretrained backbones unsuited to on-device use. We introduce RSBdSL38, 10,874 expert-validated images spanning all 38 BdSL hand signs, representing the 51 letters of the Bangla alphabet, recorded from real signers at three special-needs schools across Bangladesh. We propose a lightweight attention based convolutional network of 298,470 parameters, built from grouped bottleneck residual blocks, channel and spatial attention, a multi-scale depthwise hand-feature block, dual pooling, and Swish activations. Trained from scratch, it attains 96.37% accuracy (95.72% +- 0.54% over five seeds), within 1.08 percentage points of the best of nine ImageNet-pretrained efficient architectures under an identical protocol, using 8.5 to 68x fewer parameters and 1.3 to 21.7x fewer MACs. Retrained, it reaches 92.95 to 98.33% on six public BdSL benchmarks, 97.04% on a merged corpus, and 76.25% zero-shot on BdSL-38. Removing any architectural stage costs 7.61 to 89.30 points, against at most 3.17 for the training recipe. Grad-CAM with deletion-insertion and weight-randomization checks confirms that predictions follow the signing hand. A signer-independent split holding out 6 of 36 signers yields 85.18%. Quantized to 0.48 MB, it runs at 3.98 ms per image within a 15.5 MB footprint on a commodity smartphone. Together, RSBdSL38 and our from-scratch model turn benchmark accuracy into deployable accessibility at a fraction of pretrained-backbone cost; dataset, code, and models are released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑