发表机构
Queen Mary University of London; Jilin University; King’s College London; University College London(伦敦玛丽女王大学; 吉林大学; 伦敦国王学院; 伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对表意文字数据集不平衡问题,提出结合视觉与上下文语义的多模态学习方法,通过语义对比预训练提升字符识别性能,在多数据集及下游任务上优于现有最优方法。
AI 中文摘要
当前基于深度学习的字符视觉研究,如文本识别、字符图像去噪和历史文本补全,为学习、管理和利用字符资源提供了新方案。然而这些研究的性能仅在大规模且平衡的数据集上达到峰值,而现实中的字符数据集很少满足这一条件,尤其是表意文字语言(如中文)的数据集。由于字符使用频率的差异以及新字符的不断产生,表意文字的数据分布不平衡是常见问题。本文提出一种用于表意文字识别的新方法,引入了结合字符视觉语义和上下文语义的多模态学习方法。设计了一种新的预训练策略,通过从对应语言模型中提取每个字符的上下文语义,增强深度视觉表示,尤其针对存在不平衡和稀有实例问题的数据集。我们在多个数据集上进行实验以评估所提出的字符识别方法,并通过多个下游任务进一步验证对比预训练策略。实验结果表明,与现有最优方法相比,本文方法具有优越性。
英文摘要
Current deep learning-based character vision studies, e.g., text recognition, character image denoising, and historical text completion, are offering new solutions for learning, managing, and utilizing character resources. However, the performance of these studies peaks only with large and balanced datasets, which is a rarity with real-world character datasets, especially for logographic character languages, e.g., Chinese. The imbalance in data distribution of logographic characters is a common issue due to differences in character usage frequency and new characters being continuously created. In this paper, we propose a novel method for logographic character recognition, which introduces a multi-modal learning approach using visual semantics and contextual semantics of characters. A novel pre-training strategy is designed to enhance deep visual representations, especially for datasets suffering from issues of imbalanced and rare instances, by extracting the contextual semantics of each character from the corresponding language models. We conduct experiments across various datasets to evaluate our character recognition method and further validate the contrastive pre-training strategy by several downstream tasks. Experimental results demonstrate the superiority of our method compared to state-of-the-art methods.
CommentsACM MM 2026 paper