跨域少样本书写者自适应用于真实世界手写数学表达式识别
Cross-Domain Few-Shot Writer Adaptation for Real-World Handwritten Mathematical Expression Recognition
AI总结:
本研究针对手写数学表达式识别中训练与真实图像域差距及个人书写风格适应问题,提出书写者自适应微调流程,实验显示在多数子集上提升识别率,验证了该方法的可行性。
AI中文摘要:
手写数学表达式识别(HMER)是指将手写数学内容识别并转换为可解析的标记语言(通常是LaTeX)的任务。目前,没有任何达到最先进水平且具有竞争力的系统能够适应特定个人的书写方式,而且对于这一特定问题,训练图像(通常是数字化的或完美二值化的)与推理时用相机实际拍摄的图像之间的域差距受到的关注相当少。我们通过图像的碎片化和笔画宽度分析来表征这一域差距,并引入一种书写者自适应的微调流程到MFH-CoMER中,试图解决这一问题。我们进一步引入了作者特定的样本数据集,包含来自两位作者的五个手写类别子集,并使用McNemar检验和针对有限数据调整的置换检验进行评估。结果表明,对于数字手写,表达式识别率有所提高,字符错误率(CER)有所下降,但对于物理手写,结果则更加多样,模型在处理基础模型已经能够很好评估的手写文章时存在困难。尽管如此,我们发现自适应模型在五个子集中的四个上取得了改进的结果。统计检验结果表明,在测试的两个作者特定子集中,改进具有一致性。我们的结果指向了书写者自适应在HMER任务中的潜在可行性。
英文摘要:
Handwritten mathematical expression recognition (HMER) refers to the task of recognizing and converting handwritten mathematics into a parsable markup language, usually LaTeX. No current state-of-the-art-competitive system adjusts to the way a specific person writes, and the domain gap between training images (usually digital or perfectly binarized) and images physically taken with a camera used in inference has received fairly little attention for this specific problem. We characterize this domain gap through fragmentation and stroke-width analyses of the images as well as introduce a writer-adaptive fine-tuning pipeline to MFH-CoMER in an attempt to address it. We further introduce sample author-specific datasets, consisting of five handwriting category subsets from two authors, and evaluate using a McNemar's test and permutation tests adapted to limited data. Results suggest an increase in expression recognition rate and a decrease in CER for digital handwriting but more varied results for physical handwriting, with the model struggling for handwriting articles that the base model can already evaluate well. Nonetheless, adapted models were found to have improved results for four out of five subsets. Statistical testing results suggest a consistency in improvement for two of the tested author-specific subsets. Our results point towards the potential feasibility of writer adaptation for the HMER task.