arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31681cs.CV

Devanagari手写字符识别使用TrOCR:一种基于Transformer的模型与实时Web部署

Devanagari Handwritten Character Recognition Using TrOCR: A Transformer-Based Model with Real-Time Web Deployment

Amrit Baskota, Samyam Budhathoki, Shubham Ghimire, Abiskar Ghimire, Sarwesh Phuyal, Baskaran P

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过微调TrOCR模型实现天城文手写字符识别,达到96.05%准确率,并开发了低延迟Web应用,优于CNN模型。

中文摘要 AI 辅助

天城文是印度次大陆的古老语言之一,包含36个元音、14个辅音和10个数字。由于天城文脚本的高度复杂性,手写天城文字符的准确识别具有挑战性。本文提出了一种微调预训练TrOCR模型的方法,以准确识别天城文手写字符。该方法包括预处理机制,其中输入图像被标准化为RGB格式,分批进行标记化,并与Hugging Face数据集集成。预训练的microsoft/trocr-base-handwritten模型在近5000个字符图像的数据集上进行微调,这些图像按8:1:1的比例均匀划分为训练、评估和测试集。训练过程中通过混合精度训练、梯度检查点和早停机制进一步优化。该模型实现了3.95%的字符错误率(CER)和96.05%的字符级准确率,优于之前的基于CNN的模型。使用此http URL、Golang和FastAPI开发了一个可扩展的Web应用程序,实际部署了OCR模型,并以每请求小于5秒的延迟提供字符识别任务。本研究展示了使用TrOCR模型构建可扩展的手写天城文字符识别系统,并为未来使用Transformer进行天城文脚本识别的研究奠定了基础。

英文摘要

Devnagari is a one of the ancient language of the Indian subcontinent consisting of 36 vowels, 14 consonants and 10 numerals. The accurate recognition of handwritten Devnagari characters is challenging due to high complexity of Devnagari scripts. This paper presents a method to fine tune the pre trained TrOCR model to accurately recognize Devanagari handwritten characters. The Methodology consists of a preprocessing mechanism where input images are standardized into RGB format, tokenized in batches and integrated with Hugging Face Dataset. The pre-trained microsoft/trocr-base-handwritten model is fine-tuned over a dataset of nearly 5000 character images that are uniformly partitioned in the ratio 8:1:1 for training, evaluation and testing. Further optimization is done is the training process through mixed precision training, gradient checkpointing, and early stopping mechanism. The model achieves a character error rate (CER) of 3.95% and a character-level accuracy of 96.05%, outperforming the previous CNN based models. A scalable web application is developed using Vite.js, Golang, and FastAPI which practically deploys the OCR model and serves character recognition task with a latency less than 5 seconds per request. This study demonstrates the use of TrOCR model to build a scalable handwritten Devnagari character recognition system and also a foundation to future research on Devnagari Script Recognition using Transformers.

发表机构

  • Vellore Institute of Technology(韦洛尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑