arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30213cs.CVcs.CL

高棉语文本识别与分词联合模型研究

Towards a Joint Khmer Text Recognition and Word Segmentation

Marry Kong, Rina Buoy, Sovisal Chenda, Nguonly Taing, Masakazu Iwamura, Koichi Kise

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种采用CTC解码器的统一模型,实现高棉语文本识别与分词联合,可在不同模态基准数据集上同时完成字符识别与词边界定位,无需额外分词步骤。

中文摘要 AI 辅助

文本识别(即从文档图像中提取电子文本)对于检索增强生成(RAG)等知识检索任务至关重要。对于高棉语而言,提取的文本需额外执行分词步骤,因为高棉语不使用任何可见的单词分隔符来标识词边界。因此,高棉语“先识别后分词”的流水线需要两个独立的串行模型,这不仅易出错,还会增加大规模文档处理的延迟。本文提出一种新型高棉语文本识别与分词联合框架,该框架采用统一模型,使用连接时序分类(CTC)解码器实现快速并行解码,可分别执行带分词(b=1)和不带分词(b=0)的高棉语文本识别。在文档、场景、手写图像等不同模态的基准数据集上的实验表明,该模型不仅能识别文档图像中的字符,还能定位词边界,消除了传统串行流水线中额外的分词步骤。

英文摘要

Text recognition, or extracting electronic text from document images, has been indispensable for knowledge retrieval tasks, such as retrieval-augmented generation (RAG). For Khmer, extracted text is subject to an extra word segmentation step, as Khmer does not use any visible word delimiters to denote word boundaries. Thus, a recognition-then-segmentation pipeline for Khmer requires two separate sequential models; this is not only error-prone but also adds significant latency for large-scale document processing. This paper proposes a novel joint Khmer text recognition and word segmentation framework in a unified model. The proposed model, using a connectionist-temporal-classification (CTC) decoder for fast, parallel decoding, can be instructed to recognize Khmer text with ($b=1$) and without ($b=0$) word segmentation. Experimental results on different benchmark datasets of different document modalities (document, scene, and handwritten images) show that the proposed model can not only recognize characters in document images but also locate word boundaries, removing the need for an extra word segmentation step in a conventional sequential pipeline.

发表机构

  • Techo Startup Center(德科创业中心)
  • Ministry of Economy and Finance(经济与财政部)
  • Osaka Metropolitan University(大阪公立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑