arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35473cs.CVcs.LG

手写文本识别存在于高像素方差子空间中

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Carlos Garrido-Munoz, Jorge Calvo-Zaragoza

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明手写文本识别中判别信号集中于高像素方差子空间,因此像素重建自监督方法优于对比方法,并在多语言基准上实现最优性能。

中文摘要 AI 辅助

在手写文本识别(HTR)的自监督预训练中,像素重建方法优于对比方法,这与自然图像分类不同。我们认为这种差异源于判别信号在像素空间中的位置:对于HTR,判别信号集中在高方差方向,而低方差方向几乎没有判别信号。这预示着保留高方差像素内容的目标函数将具有最佳的迁移效果。我们在匹配的编码器、数据和评估协议下,测试了来自三个族(像素接地MIM、JEPA和对比)的六种SSL方法,在五种语言的六个手写基准上进行了测试。在完整标签下,像素接地SSL在每个基准和两种冻结探针上实现了最低的CER,暴露了其他族仅通过读出才能恢复的逐位置字符信息,并且是唯一从真实手写预训练中受益的族。像素接地表示也具有更高的标签效率。跨数据集,编码器与高方差像素子空间的对齐在每个方法内预测CER。使用预训练的LLM解码器,冻结的像素接地编码器与完全微调的有监督基线竞争;完全微调实现了最低的平均CER,并在每个基准上排名第一或第二。这些结果表明,像素重建的价值取决于判别信号在输入中的位置。

英文摘要

In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.

发表机构

  • University of Alicante(阿利坎特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑