arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24058cs.CV

全合一多语言场景文本识别:基于脚本感知的混合专家模型

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

  • Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究院)
  • Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身智能重点实验室)
  • WeChat Vision, Tencent Inc.(腾讯微信视觉团队)
  • South China University of Technology(华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen

AI总结:

针对多语言场景文本识别,提出ScriptMoE架构,利用脚本感知的混合专家和合成数据集TextMuSS-10M,在多个基准上超越现有方法,实现更轻量且更准确的统一识别器。

AI中文摘要:

多语言场景文本识别(STR)因大多数语言训练数据稀缺以及单一模型难以服务多种脚本而仍具挑战性。现有解决方案要么为每种语言部署一个识别器,导致成本增加和错误累积,要么依赖大规模视觉-语言模型(VLM),这些模型成本高昂且在许多脚本上仍不准确。在这项工作中,我们追求一种全合一的多语言识别器,它比逐语言专家更简单,比VLM更轻量,且比两者都更准确。首先,我们构建了TextMuSS-10M,一个覆盖10种脚本和229种语言的大规模合成场景文本数据集。它在真实数据不可用之处提供了平衡且充分的监督。其次,我们提出了ScriptMoE,一种脚本感知的混合专家(MoE)架构。它共享一个视觉编码器,并用稀疏MoE块替换密集解码器,该块由一个图像级路由器将每个图像分派到前两个脚本对齐的专家和一个共享专家以吸收跨脚本知识。在我们组装的TextMuSS-Bench(10种脚本,10,899张图像)上的广泛实验表明,ScriptMoE达到了82.06%的最高准确率,比最强的STR基线高出1.31%。在CC-OCR端到端多语言任务中,仅将PP-OCRv5中的识别器替换为ScriptMoE,就将F1分数从65.71%提升至80.89%,以一小部分参数数量略微超越了最佳VLM(80.73%)。

英文摘要:

Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.

补充信息

相关深度报道

↑