arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37195cs.CV

探索手写文本识别中的上下文学习

Exploring In-Context Learning for Handwritten Text Recognition

Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza

首次发表
浏览论文内容

中文总结 AI 辅助

本文探索利用预训练视觉-语言模型进行上下文学习,在不更新参数的情况下构建手写文本识别流程,并在域内和跨域场景中验证其性能与传统HTR相当,同时指出采样方法需进一步优化。

中文摘要 AI 辅助

手写文本识别(HTR)系统已成为历史文献数字化不可或缺的工具。它们不仅降低了时间和成本,还通过生成转录文本,使人们能够民主地访问和处理其内容。然而,当前HTR领域的文献主要集中于需要大量标注样本才能达到满意性能的专用模型。我们探索了使用预训练的视觉-语言模型(VLM)进行上下文学习(In-Context Learning),以在不更新模型参数的情况下构建转录流程。随后,我们在多个数据集和模型上评估了这一流程,并证明通用VLM可以被有效地教会如何从图像中转录手写文本。为了评估我们的观察结果如何转化为实际应用,我们在跨域(CD)场景中评估了性能,其中上下文示例取自与查询图像不同的数据集。在受控的域内(ID)场景和现实的CD场景中,结果遵循相同的模式。首先,随着上下文大小的增长,错误范围预计会向平均性能收窄。因此,较大的上下文大小会牺牲基于oracle最佳采样的性能,以换取更低的预期错误率。结果表明,在没有任何参数更新的情况下,该方法在存在域偏移时具有很强的与传统HTR竞争的潜力。此外,我们展示并论证了某些上下文采样方法比其他方法效果更好,并建议未来的研究应投入更多努力来寻找理想的采样方法。

英文摘要

Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses mostly on specialized models that require large amounts of annotated samples to achieve satisfactory performance. We explore the use of In-Context Learning with pre-trained Vision-Language Models (VLMs) to create a transcription pipeline without updating the model's parameters. We then evaluate this pipeline across multiple collections and models, and demonstrate that general-purpose VLMs can be effectively taught how to transcribe handwritten text from images. To assess how our observations may translate to practical applications, we evaluate the performance in a Cross-Domain (CD) scenario, where context examples are drawn from a different collection than the query image. Results in both the controlled In-Domain (ID) scenario and the realistic CD scenario follow the same patterns. First, as context size grows, the error range is expected to narrow towards the average performance. Thus, larger context sizes sacrifice the performance of the oracle-best sampling for lower expected error rates. The results obtained show that, without any parameter updates, this methodology has strong potential to compete with traditional HTR in the presence of domain shift. Moreover, we show and argue that some context samplings work better than others and suggest more effort should be put into finding an ideal sampling method in future work.

发表机构

  • University of Alicante(阿利坎特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑