发表机构
Faculty of Electrical Engineering, Universiti Teknologi Malaysia; Faculty of Artificial Intelligence, Universiti Teknologi Malaysia; Department of Electrical and Electronic Engineering, American International University-Bangladesh; Department of Electrical Engineering, Balochistan University of Information Technology, Engineering and Management Sciences(马来西亚工艺大学电气工程学院; 马来西亚工艺大学人工智能学院; 孟加拉国美国国际大学电气与电子工程系; 巴基斯坦俾路支省信息技术、工程与管理科学大学电气工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于PRISMA指南,通过分析2015-2025年97项研究,系统综述面向文本识别的机器学习模型及应用,梳理其演进、挑战并提出优化建议,为OCR领域研究提供基础。
AI 中文摘要
基于优选报告项目用于系统评价和荟萃分析(PRISMA)指南,本文献综述对过去十年的光学字符识别(OCR)AI模型演进展开广泛评估,涵盖2015年1月至2025年1月间筛选的97项研究,分析其模型转变、应用领域、数据类型、语言覆盖及挑战。传统OCR模型难以应对脚本变体、书写风格及退化文档,而机器学习驱动的OCR在异构文本数据处理上已显著进步;本研究识别关键OCR模型,分析其性能、优势与局限,发现OCR技术已演进至可处理结构化与非结构化文本、场景文本识别及多语言处理。未解决的挑战包括代表性不足语言资源有限、手写文本高变异性、字符间视觉相似性及实时OCR应用的约束。针对这些问题,提出自监督学习、多模态AI、自动机器学习(AutoML)、AI辅助后处理、微型机器学习(TinyML)及创建脚本匹配联合语料库等有前景的方法,未来建议旨在提升OCR精度,应对实时工业应用的挑战,该研究将为OCR领域的未来研究提供指导并奠定基础。
英文摘要
Optical Character Recognition (OCR) for text recognition using machine vision has significantly improved, particularly when handling heterogeneous textual data. Traditional OCR models struggle with script variations, writing styles, and degraded documents. Advancements in technology are leading to new AI models with improved architecture for handling multiple languages and complex data formats. Despite this progress, a comprehensive evaluation of OCR advancements remains limited. Based on the established preferred reporting items for systematic reviews and meta-analysis (PRISMA) guidelines, this literature review presents an extensive assessment of OCR research to trace the evolution of AI models over the past decade. It explores the transition in AI models, application domains, data types, linguistic coverage, and challenges. Through a detailed analysis of 97 selected studies published during January 2015 - January 2025, key OCR models are identified, and their performance, strengths, and limitations are analyzed. The findings highlight how OCR technologies have evolved to address structured and unstructured text, scene text recognition, and multilingual processing. Unresolved challenges include limited resources for underrepresented languages, high variability in handwritten text, visual similarity among characters, and constraints in real-time OCR applications. To address these issues, several promising approaches are proposed. Key suggestions include self-supervised learning, multimodal AI, automated machine learning (AutoML), AI-assisted postprocessing, tiny machine learning (TinyML), and the creation of joint corpora for script matching. The future recommendations aim to enhance OCR accuracy and tackle the challenges identified for real-time industrial applications. This study will guide future research and establish a foundation for OCR field.
CommentsPublished in IEEE Access, 2025. 24 pages, 16 figures, 5 tables
Journal refvol. 13, pp. 177647-177670, 2025
DOI:10.1109/ACCESS.2025.3618109