arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28195cs.CV

UniLipi:面向印度历史手稿的统一多脚本光学字符识别系统

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla

首次发表
浏览论文内容

中文总结 AI 辅助

UniLipi是面向印度历史手稿的统一多脚本OCR模型,可处理复杂手稿场景,利用合成数据减少真实标注依赖,还可作为基础模型适配多类脚本,支持手稿编目。

中文摘要 AI 辅助

针对手写印度手稿的光学字符识别(OCR)是大规模数字化及实现手稿遗产计算访问的关键。然而现有方法通常针对单一脚本开发,需要大量特定脚本的定制,这限制了其在多样藏品中的可扩展性与实际部署。本文提出UniLipi,一个在单一框架内联合13种印度脚本训练的手写印度手稿统一多脚本OCR模型。UniLipi可直接处理真实手稿场景,包括线条几何的极端变化、线条长度的大幅波动,以及孔洞、污渍、图画插图等非文本手稿实体造成的部分中断。为在超低资源条件下有效运行,该模型利用脚本感知的合成手稿数据生成,大幅减少了对大量真实标注数据的依赖。除历史手稿外,UniLipi还可作为有效的预训练基础模型,其学习到的表示不仅能在当代印度手写体上实现良好的OCR性能,还可扩展至藏文、意大利文、拉丁文、中文等多种非印度脚本。除转录外,UniLipi还能预测脚本身份和每行原生字符数,支持实际的手稿编目工作流程。

英文摘要

Optical character recognition (OCR) for handwritten Indic manuscripts is essential for large-scale digitization and computational access to manuscript heritage. However, existing approaches are typically developed for one script at a time and require substantial script-specific customization. This limits scalability and practical deployment across diverse collections. We present UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework. UniLipi directly handles realistic manuscript conditions, including extreme variation in line geometry, large variation in line length, and partial interruptions caused by non-textual manuscript entities such as holes, stains, or pictorial illustrations. To operate effectively under ultra low-resource conditions, the model leverages script-aware synthetic manuscript data generation, substantially reducing reliance on large volumes of real annotated data. Beyond historical manuscripts, we show that UniLipi serves as an effective foundational pretrained model. Specifically, its learned representations enable good OCR performance for contemporary Indic handwriting and extend to several non-Indic scripts, including Tibetan, Italian, Latin, and Chinese scripts. In addition to transcription, UniLipi predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.

发表机构

  • International Institute of Information Technology Hyderabad(国际信息技术研究所海得拉巴分校)
  • Gitam(吉塔姆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑