arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExpertHTR:基于多任务学习与稀疏混合专家模型的统一手写文本识别

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy

arXiv 2609.12705首次发表:更新:

发表机构

University of Information Technology, Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam; AJ Technologies, Nagoya, Japan; RIKEN Center for Advanced Intelligence Project, Tokyo, Japan(胡志明市信息技术大学; 越南国立大学胡志明市分校; AJ科技公司; 理化学研究所先进智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对手写文本识别资源分散、联合训练困难的问题,提出统一框架ExpertHTR,利用互补监督和稀疏混合专家模型,在七个基准上多数超越通用系统并刷新IAM段落级纪录。

AI 中文摘要

手写文本识别资源通常规模较小,且分散在不同语言、文字、文档结构和标注格式的多个数据集中,这使得联合的页面级训练变得困难。我们提出ExpertHTR,一个统一的视觉-语言框架,通过互补监督和条件模型容量来解决这一问题。来自异构数据集的结构化标注首先通过通用的页面-区域-行表示进行组织,并用于构建四个相关的训练任务:完整转录、物理行覆盖、文本定位和局部识别,无需额外的人工标注。在联合训练的密集模型基础上,ExpertHTR引入了稀疏混合专家架构,包含一个始终激活的共享分支和条件路由的全MLP专家。Sparsegen允许激活的路由专家数量随隐藏表示而变化,而路由正则化减少了在少数专家上的持续集中。在七个异构手写基准上的实验表明,互补监督持续改善了仅使用页面转录的训练效果,而联合多源训练在大多数数据集上带来了进一步的提升。所提出的稀疏专家模型在七个数据源中的六个上进一步优于密集基线。最终的统一模型在大多数基准上大幅优于所评估的通用OCR和视觉-语言系统,并在IAM段落级基准上达到了最先进的性能,而专门的手写识别系统在几个具有挑战性的数据集上仍然更强。

英文摘要

Handwritten text recognition resources are often small and distributed across collections that differ in language, script, document structure, and annotation format, making joint page-level training difficult. We propose ExpertHTR, a unified vision-language framework that addresses this problem through complementary supervision and conditional model capacity. Structural annotations from heterogeneous datasets are first organized through a common Page-Region-Line representation and used to construct four related training tasks for complete transcription, physical-line coverage, text localization, and localized recognition, without requiring additional manual labels. Building on a jointly trained dense model, ExpertHTR introduces a sparse Mixture-of-Experts architecture with an always-active shared branch and conditionally routed full-MLP experts. Sparsegen allows the number of active routed experts to vary with the hidden representation, while routing regularization reduces persistent concentration on a small subset of experts. Experiments on seven heterogeneous handwriting benchmarks show that complementary supervision consistently improves training with page transcription alone, while joint multi-source training provides further gains on most datasets. The proposed sparse expert model further improves the dense baseline on six of the seven sources. The final unified model also substantially outperforms the evaluated general-purpose OCR and vision-language systems on most benchmarks and achieves state-of-the-art performance on the IAM paragraph-level benchmark, while specialized HTR systems remain stronger on several challenging collections.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑