发表机构
Indian Institute of Technology Gandhinagar(印度理工学院甘地讷格尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出InscriptionOCR,一个端到端AI框架,涵盖图像增强、OCR、音译和神经翻译,并发布含20万字符图像和2000句对的数据集,以解决阿育王婆罗米文铭文理解难题。
AI 中文摘要
古代文字图像修复是计算机视觉中的一个基本问题,因为它直接影响对历史文献和铭文的可靠分析与解读。阿育王婆罗米文是一种古代文字,在公元前3世纪阿育王统治时期广泛使用,主要用于普拉克里特语的铭文。这些铭文,包括大、小岩石敕令和石柱敕令,构成了一个有价值但很大程度上尚未被开发的计算分析数据来源。铭文图像的退化性质以及缺乏标准化的数字资源,给自动化处理带来了重大挑战。我们提出了一个用于理解古代铭文的端到端基于人工智能的框架,涵盖图像增强、光学字符识别(OCR)、音译和神经机器翻译(NMT)。所提出的流程处理直接从石碑铭文拍摄的低质量图像,执行图像修复和婆罗米文字符识别,将识别出的字符映射到罗马字母,最后将生成的普拉克里特语文本翻译成英语。我们还引入了两个新数据集:(i)InscriptionOCR数据集:迄今为止最大的公开可用的婆罗米文数字OCR数据集,包含约600个类别中的超过200,000个字符图像,以及(ii)一个双语普拉克里特语-英语平行语料库,包含超过2,000个句子对,用于NMT。我们相信,所提出的框架和数据集将促进古代文字分析、低资源OCR和数字金石学领域的未来研究。
英文摘要
Ancient script image restoration is a fundamental problem in computer vision, as it directly affects the reliable analysis and interpretation of historical documents and inscriptions. Ashokan Brahmi is an ancient script extensively used during the reign of Emperor Ashoka in the 3rd century BC, primarily for inscriptions in Prakrit. These inscriptions, including major and minor rock and pillar edicts, constitute a valuable yet largely unexplored source of data for computational analysis. The degraded nature of inscription imagery and the lack of standardized digital resources pose significant challenges for automated processing. We present an end-to-end AI-based framework for understanding ancient inscriptions that encompasses image enhancement, optical character recognition (OCR), transliteration, and neural machine translation (NMT). The proposed pipeline processes low-quality images captured directly from stone inscriptions, performs image restoration and Brahmi script character recognition, maps the recognized characters to the Roman script, and finally translates the resulting Prakrit text into English. We also introduce two new datasets: (i) InscriptionOCR Dataset: the largest publicly usable digital OCR dataset for Brahmi script to date, consisting of over 200,000 character images across about 600 classes, and (ii) a bilingual Prakrit-English parallel corpus comprising over 2,000 sentence pairs for NMT. We believe that the proposed framework and datasets will facilitate future research in ancient script analysis, low-resource OCR, and digital epigraphy.