arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09449cs.CV

从全局对齐到局部锚定:基于部首验证的零样本汉字识别

From Global Alignment to Local Grounding: Zero-Shot Chinese Character Recognition with Radical Verification

Yu-Heng Shih, Bing-Chen Wu, Tsz-To Wong, Ting-En Yen, Hong-Han Shuai, Bin-Hua Hsieh, Chien-An Chen, Yi-Ren Yeh, Ching-Chun Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对零样本汉字识别中全局匹配忽略部首空间布局导致泛化差的问题,提出全局到局部两阶段框架:STG-CLIP注入几何先验实现高召回检索,RVM模块验证局部部首匹配,在ICDAR2013上以83.06%的top-1准确率取得最优性能。

中文摘要 AI 辅助

零样本汉字识别(ZS-CCR)旨在识别训练过程中从未见过的类别字符,通常依赖于已见字符与未见字符之间共享的组成结构。近期的CLIP风格方法利用汉字结构描述码(IDS)表示这种结构,并在共享嵌入空间中将IDS与字形图像对齐。然而,这些方法依赖于单一的全局图像-IDS相似度,该相似度忽略了部首的空间布局,并且仅从已见类别中隐式学习,导致对未见类别的泛化能力较差;此外,全局匹配通常能在候选集中检索到正确字符,但当字符仅在细微的局部部首上存在差异时,却无法将其排到首位。为解决这些问题,我们提出了一种从全局到局部的两阶段框架。在第一阶段,STG-CLIP为IDS添加了显式的树位置和部首级几何先验,生成空间感知原型,该原型在已见和未见类别间提供一致的空间描述,以实现高召回率的全局检索。在第二阶段,部首验证模块(RVM)以每个检索候选字符的部首实例作为查询,验证相应部首是否能与输入字形中的空间兼容区域匹配。基于边距的门控规则仅在领先的全局候选获得相似相似度分数时激活RVM。在ICDAR2013基准上的实验表明,我们的方法在字符级零样本设置下达到了最先进的性能,在2,755个已见类别上获得了83.06%的top-1准确率。消融研究进一步表明,显式几何先验和部首级验证提供了互补性的改进。

英文摘要

Zero-shot Chinese character recognition (ZS-CCR) aims to recognize characters whose categories are never observed during training, and typically relies on the compositional structure shared between seen and unseen characters. Recent CLIP-style methods represent this structure with the Ideographic Description Sequence (IDS) and align it with glyph images in a shared embedding space. However, they rely on a single global image--IDS similarity that discards the spatial layout of radicals and, being learned only implicitly from seen classes, generalizes poorly to unseen ones; moreover, global matching often retrieves the correct character within the top candidates yet fails to rank it first when characters differ only in subtle local radicals. To address these issues, we propose a global-to-local two-stage framework. In the first stage, STG-CLIP augments the IDS with explicit tree-position and radical-level geometric priors, yielding a spatial-aware prototype that provides a consistent spatial description across seen and unseen categories for high-recall global retrieval. In the second stage, the Radical Verification Module (RVM) uses the radical instances of each retrieved candidate as queries to verify whether the corresponding radicals can be matched to spatially compatible regions in the input glyph. A margin-based gating rule activates the RVM only when the leading global candidates receive similar similarity scores. Experiments on the ICDAR2013 benchmark demonstrate that our method achieves state-of-the-art performance under the character-level zero-shot setting, obtaining 83.06% top-1 accuracy with 2,755 seen classes. Ablation studies further show that the explicit geometric priors and radical-level verification provide complementary improvements.

发表机构

  • National Yang Ming Chiao Tung University(国立阳明交通大学)
  • E.SUN Financial Holding Co., Ltd.(玉山金融控股股份有限公司)
  • National Kaohsiung Normal University(国立高雄师范大学)

机构由 AI 辅助整理,请以论文原文为准。

↑