发表机构
College of Computer Science, Chengdu University; Key Laboratory of Digital Innovation of Tianfu Culture, Sichuan Provincial Department of Culture and Tourism, Chengdu University; Machine Intelligence Laboratory, College of Computer Science, Sichuan University; Tianfu Jincheng Laboratory; Network Information Centre, Hainan College of Economics and Business(成都大学计算机科学学院; 成都大学四川省文化和旅游厅天府文化数字创新重点实验室; 四川大学计算机科学学院机器智能实验室; 天府锦城实验室; 海南经贸职业技术学院网络信息中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
为解决手写和印刷文本分割中深度学习模型计算成本高的问题,提出轻量级框架。先通过句子级连通分量分割算法提取片段,再设计区域感知手写描述符,简单分类器可集成,实验表明该框架在准确率和效率上优于现有方法,适合轻量级部署。
AI 中文摘要
随着教育和办公环境中对纸质文档再利用需求的增加,手写和印刷文本的准确分割成为文档数字化的关键步骤。尽管已开发出众多深度学习模型,但高计算成本限制了在资源受限边缘设备上的部署。为此,我们提出一个针对计算能力严重受限设备优化的轻量级框架。该方法始于句子级连通分量分割算法以提取文档图像中的连贯句子级片段,接着设计新颖的区域感知手写描述符(RHD)在句子级别捕获人类手写的内在变异性,简单传统分类器可与设计的描述符无缝集成,在区分手写和印刷句子级文本图像上展现出强大分类性能。在自建的多语言高质量手写和印刷文本分割注释数据集(MAD-HPTS)及公共基准PHD-AS上进行大量实验,结果表明所提框架在准确性和计算效率上均优于当前最先进方法。在MAD-HPTS上,与领先的深度神经网络基线相比,我们的方法仅牺牲1.4%的准确率,但推理速度提高了8倍多,非常适合轻量级部署。
英文摘要
With the increasing demand for reusing paper documents in educational and office settings, accurate segmentation of handwritten and printed text has become a crucial step in document digitization. Although numerous deep learning models have been developed for this task, their high computational cost limits deployment on resource-constrained edge devices. To address this challenge, we present a lightweight framework optimized for efficient performance on devices with severely limited computational capacity. Our approach begins with the Sentence-level Connected Component Segmentation algorithm, aimed at extracting coherent sentence-level segments from document images. We then design a novel Region-aware Handwriting Descriptor (RHD) to capture the intrinsic variability of human handwriting at the sentence level. A simple conventional classifier can then be seamlessly integrated with our designed descriptor, demonstrating strong classification performance for distinguishing handwritten and printed sentence-level text images, highlighting that the proposed descriptor is agnostic to the choice of classifier. Extensive experiments are performed on our self-constructed Multilingual High-Quality Annotated Dataset for Handwritten and Printed Text Segmentation (MAD-HPTS) and a public benchmark PHD-AS, and the experimental results demonstrate that the proposed framework outperforms current state-of-the-art methods in both accuracy and computational efficiency. On MAD-HPTS, our method sacrifices only 1.4% accuracy compared to the leading deep neural network baseline, yet achieves more than 8 times speedup in inference, making it well-suited for lightweight deployment.
Comments21 pages, 8 figures, and 5 tables