发表机构
Carnegie Mellon University Africa(卡内基梅隆大学非洲分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CASL手语识别的低准确率问题,提出轻量级Transformer模型TransSLR,在CASL-W60基准上取得80.39%的最新准确率,且计算开销低,适用于资源受限环境。
AI 中文摘要
针对代表性不足语言的自动手语识别仍是一个未解决的问题。中非手语(CASL)就是这一差距的典型例子:唯一可用的基准CASL-W60报告的最佳准确率为69.93%,且我们发现,对高资源模型进行微调的通用启发式方法无法缩小这一差距。这一失败源于两个叠加因素:可用CASL数据的规模有限,以及CASL与WLASL等大规模语料库之间存在显著的词汇和视觉领域差距,这使得预训练表示基本无用。为解决该问题,我们提出TransSLR,一种轻量级时间Transformer编码器,在64帧归一化姿态序列上从头开始训练,采用平均池化和分类头。通过基于几何关键点表示而非原始RGB操作,TransSLR无需依赖视觉外观即可实现独立于手语使用者的泛化。在CASL-W60基准上,TransSLR达到了80.39%的最新准确率,比之前的最佳结果高出10.46%。除准确率外,我们的仅编码器设计显著降低了计算开销,使其可在资源受限环境中部署。我们在CASL-W60基准上进行了广泛实验,与基于RGB和多模态基线进行比较,证明TransSLR达到了最新性能。
英文摘要
Automated sign language recognition for underrepresented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this challenge: the CASL-W60 benchmark contains 60 isolated signs, and the best published result is 69.93%. We present TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on normalized 64-frame pose sequences. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves 80.39% top-1 accuracy and 91.07% top-5 accuracy on a signer-independent evaluation split. This result is 10.46 percentage points higher than the published 69.93% result. TransSLR has 8.67 million trainable parameters. We also evaluate zero-shot transfer from a high-resource sign-language model, which achieves 0.00% exact-match accuracy under our manual gloss-matching protocol. Our experiments compare pose-based, RGB-based, and multimodal approaches to sign-language recognition for a low-resource language.
CommentsThis paper has been accepted for oral presentation at Deep Learning Indaba 2026 which will be hosted on the IJCAI hosting platform. paper link: https://chairingtool.com/conferences/dli2026/main-track/submissions/356