arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于原始和合成深度图像的点云数据模型的手语识别

Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models

Rustem Ozakar, Eyup Gedikli

arXiv 2608.09400首次发表:更新:

发表机构

Erzurum Technical University; Trabzon University(埃尔祖鲁姆技术大学; 特拉布宗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究利用Depth Anything V2生成合成深度图像,结合三个手语数据集,采用PointNet等模型对比原始与合成深度图像点云的手语识别性能,发现多数模型中原始数据性能更优,部分模型合成数据更优。

AI 中文摘要

手语识别研究大多依赖RGB图像,而提供深度图像的手语数据集有限。从深度图像获取的点云可用于PointNet等神经网络的手语识别。近年来,多种神经网络被用于从单目RGB图像生成逼真的深度图像。本研究使用Depth Anything V2网络从RGB图像创建合成深度图像,为此采用了三个同时包含RGB和深度图像的手语数据集:实时ASL手指拼写数据集、KArSL数据集、AUTSL数据集。使用多种PointNet架构,测量了从原始深度图像和合成深度图像创建的点云数据的手语识别分类准确率。从原始和合成点云中,采用基于帧、点手势图和长短期记忆数据模型进行分类,并比较了它们的性能。结果显示,原始和合成数据在大多数模型中均达到可接受的性能。总体而言,基于原始深度的点云模型性能优于合成模型,但在部分模型中,基于合成深度的模型性能优于原始模型。

英文摘要

Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Point clouds obtained from depth images can be used for sign language recognition with neural networks like PointNet. In recent years, various neural networks are used for generating realistic depth images from monocular RGB images. In this work, synthetic depth images were created from RGB images using Depth Anything V2 network. For this purpose, three sign language datasets (Real-time ASL Fingerspelling, KArSL, AUTSL) which contain both RGB and depth images were used. Classification accuracies of the point cloud data created from both original and synthetic depth images using various PointNet architectures were measured for sign language recognition. From the original and synthetic point clouds, frame based, Point Gesture Map and Long Short Term Memory data models were used for classification and their performances were compared. In the results, both original and synthetic based data achieved acceptable performance in most models. In general, original depth based point cloud models performed better than synthetic ones, however in some models synthetic depth based models performed better than the originals.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑