开发一个用于提取韩语发票信息的OCR模型
Developing an OCR model for Extracting Information from Invoices with Korean Language
浏览论文内容
中文总结 AI 辅助
本文提出一种结合深度学习与图像预处理的高效OCR模型,用于自动提取韩语发票信息,在收集的发票上达到87%的F1分数,处理时间可忽略。
中文摘要 AI 辅助
发票是包含各种信息的商业文件,包括所购物品、时间和总金额。因此,提取重要信息变得至关重要。存储的信息服务于不同的目的。韩语是大约8000万人的母语,不仅在韩国(包括南韩和北韩)发挥重要作用,而且在许多其他国家如越南、菲律宾等大量韩国公司所在的国家也扮演重要角色。在此背景下,为了自动从韩语发票中提取适当信息,我们提出了一种高效的OCR(光学字符识别)模型,该模型将深度学习模型与一些图像预处理技术相结合。所提出的OCR模型在一组丰富的收集发票上进行了评估,结果显示在可忽略的处理时间内可以达到87%的F1分数。
英文摘要
Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money. Making the extraction of important information crucial. The stored information serves different purposes. Korean language is the native language of about 80 million people, playing an important role in not only South and North Korea but also in many other countries such as Vietnam, Philippine where a large number of Korean companies are located. In this context, to automatically extract proper information from the invoices with Korean language, we propose an efficient Optical Character Recognition (OCR) model in which a deep learning model is combined with some image preprocessing techniques. The proposed OCR model is assessed in a rich set of collected invoices showing that 87% F1-score can be achieved with negligible time processing.
发表机构
- University of Engineering and Technology, Vietnam National University-Hanoi(河内国家大学工程技术大学)
- Posts and Telecommunications Institute of Technology(邮电技术学院)
机构由 AI 辅助整理,请以论文原文为准。