arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从转录的孟加拉语文本中进行开放词汇单词识别

Open Vocabulary Word Recognition From Transcribed Bangla Texts

Faias Satter, Sk. Md. Masudul Ahsan

arXiv 2610.01134首次发表:更新:

发表机构

Department of Computer Science; Engineering Khulna University of Engineering \& Technology Khulna-9203, Bangladesh; Engineering Khulna University of Engineering \& Technology Khulna-9203, Bangladesh Email: 1 3(; ; )

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用SSD、Faster R-CNN及集成模型识别手写孟加拉语单词,改进非极大值抑制,在9841图像数据集上集成模型F1达92.61%,单词识别率96.12%。

AI 中文摘要

光学字符识别(OCR)可以利用技术扫描纸张并提取文本,使人们的工作更加轻松。虽然软件行业中有各种OCR系统可用,但要找到可靠的孟加拉语等效解决方案仍需大量工作。在手写文本方面,情况则更为特殊。从单词图像中识别单词是任何OCR过程中最关键的一步,这是在从文本图像中分割单词之后的第二阶段。如果这一阶段失败,无论其他阶段表现如何,OCR的整体性能都会很差。本研究旨在利用深度学习识别手写孟加拉语单词图像中的单词。使用了三种目标检测模型:带有MobileNetV2的SSD、带有InceptionResNetV2的Faster R-CNN,以及这两种模型的集成模型,用于训练和测试手写单词图像。引入了一种改进的非极大值抑制(Non-Maximum Suppression)来增强模型结果的有效性。编制了一个包含9841张手写孟加拉语单词图像的自定义数据集,其中包含来自不同个体的多样化手写风格。所有三个模型的性能均在测试数据集上进行了检验,集成模型表现最为突出,F1分数达到92.61%。此外,在单词级别上,集成模型在一定程度上正确识别了96.12%的单词。该系统可以通过引入后处理阶段来进一步改进,以纠正系统产生的错误。

英文摘要

An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are available in the software industry, finding a reliable equivalent solution for Bangla takes much work. When it comes to handwritten texts, the situation is much more unusual. Recognizing words from word images is the most critical stage in any OCR process. It is the second stage after segmenting words from text pictures. If this stage fails, the overall performance of the OCR will be poor, regardless of how well the other phases perform. This study aims to recognize words using deep learning in a handwritten Bangla word image. Three object detection models, SSD with MobileNetV2, Faster R-CNN with InceptionResNetV2, and an ensemble model of these two, have been used to train and test handwritten word images. A modified Non-Maximum Suppression has been introduced to enhance the effectiveness of the models' results. A customized dataset of 9841 handwritten Bangla word images has been compiled, featuring diverse handwriting styles from various individuals. All three models' performances have been checked against the test dataset, and the ensemble model has been the most impressive, with an F1-score of 92.61%. Also, at the word level, the ensemble model correctly recognizes 96.12% of the words to some extent. The system can be further improved by introducing a post-processing phase to correct errors generated by the system.

Comments6 pages, 4 figures, 5 tables. Accepted version of the paper published in the 2023 26th International Conference on Computer and Information Technology (ICCIT). Code: https://github.com/FaiasPromit/Open-Vocabulary-Word-Recognition-From-Transcribed-Bangla-Texts.git

Journal ref2023 26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2023

DOI:10.1109/ICCIT60459.2023.10441393

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑