arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

车内手语语料库(ICSL):用于受限空间手语识别的多模态资源

The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition

Raviteja Boddu, Guilherme Vieira Leite, Joed Lopes da Silva, Ângelo Benetti, Isabela Barbieri, Natália de Melo Afonso, Thyago Santos, Helio Pedrini, Felipe Venâncio Barbosa, José Mario De Martino, Munir Georges, Alessandro Zimmer

arXiv 2607.11341首次发表:更新:

发表机构

Technische Hochschule Ingolstadt (THI); Universidade Estadual de Campinas (UNICAMP); Universidade de São Paulo (USP)(因戈尔施塔特技术大学; 坎皮纳斯州立大学; 圣保罗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究共享出行服务中受限空间的手语识别挑战,提出ICSL数据集,含高精度实验室数据和现实多模态车内记录,为相关比较分析提供基础,描述了语料库多方面细节,助力提升聋人社区公共交通可达性及乘客生活质量。

AI 中文摘要

本文探讨了在共享出行服务(如出租车、拼车或拼车平台)中使用手语的挑战。在现实世界的受限环境,特别是车辆内部使用手语识别(SLR)在很大程度上仍未得到探索。为推动该领域的研究,我们提出了用于巴西手语(Libras)的车内手语(ICSL)数据集,长期目标是改善聋人和听力障碍社区的公共交通可达性。该数据集包括高精度实验室运动捕捉(MoCap)数据以建立理想化语言基线,以及使用二维相机和三维飞行时间传感器捕获的现实世界多模态车内记录。数据集为合成手语虚拟动画与录制的真实手语翻译视频之间的比较分析提供了基础,有助于未来对强大的“野外”SLR模型和域适应的研究。我们详细描述了语料库的用例、设置、数据收集协议和元数据结构。总共记录了一个超过150万帧的多模态数据集,包括上述跨各种车内场景的Libras用户的同步多模态流。语料库提供了词汇手语和非词汇手语元素的词汇注释,专门用于支持受限空间识别的深度神经网络的训练和评估。车内手语提供了一个在受限、遮挡和非正面环境中的技术上重要的示例。在认识到聋人社区已经采用的各种交流策略的同时,识别汽车特定的限制为研究提高车内可达性和乘客生活质量提供了有用的垫脚石。

英文摘要

This paper addresses the challenges of using sign language within shared mobility services, such as taxis, carpools, or ride-sharing platforms. The use of sign language recognition (SLR) in real-world, confined environments, specifically vehicle interiors remains largely unexplored. To motivate research in this area, we present the In-Car Sign Language (ICSL) dataset for Brazilian Sign Language (Libras), with the long-term goal of improving public transport accessibility for the Deaf and Hard-of-Hearing community. The dataset consists of: (1) high-precision laboratory motion capture (MoCap) data to establish an idealized linguistic baseline and (2) real-world multi-modal in-car recordings captured using a 2D camera and 3D Time-of-Flight sensors. The dataset provides a basis for comparative analyses between synthesized signing avatar animations and recorded real signing interpreter videos, which enable future research into robust "in-the-wild" SLR models and domain adaptation. We describe in detail the use cases, the setup, the data collection protocol, and the metadata structure of the corpus. In total, we recorded a multimodal dataset exceeding 1.5 million frames, comprising the synchronized multimodal streams described above featuring Libras users across various in-car scenarios. The corpus is provided with gloss annotation of lexical signs and non-lexical sign language elements specially designed to support the training and evaluation of deep neural networks for constrained space recognition. In-vehicle signing offers a technically significant example of a constrained, occluded, and non-frontal environment. While recognizing the diverse communication strategies already employed by the Deaf community, identifying automotive-specific limitations provides a useful stepping stone for research into enhancing in-car accessibility and passenger quality of life.

CommentsPublished in the Proceedings of the LREC2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion Original publication: https://www.sign-lang.uni-hamburg.de/lrec/pub/26.html The paper is distributed under the CC BY-NC 4.0 license. Link to paper: https://www.sign-lang.uni-hamburg.de/lrec/pub/26033.html

Journal refProceedings of the LREC2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑