arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从转录孟加拉语段落中进行颜色无关的词分割

Color Independent Word Segmentation From Transcribed Bangla Passages

Faias Satter, Noor Masrur, Sk. Md. Masudul Ahsan

arXiv 2610.01191首次发表:更新:

发表机构

Department of Computer Science; Engineering Khulna University of Engineering \& Technology Khulna-9203, Bangladesh; Engineering Khulna University of Engineering \& Technology Khulna-9203, Bangladesh Email: 1 2 3(; ; )

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种颜色无关的手写孟加拉语文本词分割方法,在包含阴影干扰的自定义数据集上实现91.20%的F1分数,为OCR系统奠定基础。

AI 中文摘要

光学字符识别(OCR)系统可以扫描纸张并提取文本,使人们的工作更加轻松。虽然软件领域中有许多可用的OCR系统,但要找到可靠的孟加拉语等效解决方案却很困难。对于手写文本而言,这种情况更为罕见。任何OCR的第一步基础步骤是从文本图像中分割单词。如果此阶段失败,无论后续阶段表现多么有前景,整个OCR的性能都会很差。本研究旨在分割手写孟加拉语文本图像中的单词。该研究可应用于任何智能手机拍摄的图像,无论纸张和墨水的颜色和类型如何。此外,由于智能手机拍摄的图像可能产生阴影干扰,为本研究构建的自定义数据集包含了可能遇到的每一个障碍。对于7374个单词,共生成了7278个边界框,其召回率为90.60%,精确率为91.80%,F1分数为91.20%。该系统可以通过对包含多个单词的边界框进行嵌套操作,或将自适应阈值和膨胀滤波器大小调整到更精确的水平来进一步改进。

英文摘要

An optical character recognition(OCR) system can scan paper and extract text, making people's jobs easier. While numerous OCR systems are accessible in the software sector, finding a dependable equivalent solution for Bangla is tough. When it comes to handwritten texts, the case is even more rare. The first fundamental step to any OCR is to segment words from text images. If this stage fails, the total OCR's performance will be poor no matter how promising the later stages perform. This research aims to segment words in a handwritten Bangla text image. This research can be implemented on any smartphone-captured image, irrespective of the color and type of paper and ink. Furthermore, as smartphone-captured images can create shadow interferences, the custom dataset built for this research is created in such a way that every possible obstacle that can be faced is included. For 7374 words, a total of 7278 bounding boxes are generated, which have recall of 90.60%, precision of 91.80%, and F1-score of 91.20%. The system can be further improved with nested operations on bounding boxes containing several words or by adjusting the adaptive thresholding and dilation filter sizes to a more precise level.

Comments6 pages, 8 figures, 6 tables. Accepted version of the paper published in the 2023 6th International Conference on Electrical Information and Communication Technology (EICT). Code: https://github.com/FaiasPromit/Color-Independent-Word-Segmentation-From-Transcribed-Bangla-Passages.git

Journal ref2023 6th International Conference on Electrical Information and Communication Technology (EICT), 2023

DOI:10.1109/EICT61409.2023.10427730

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑