发表机构
Electrical Engineering Department, Chabahar Maritime University; Department of Systems Design Engineering, University of Waterloo(恰巴哈尔海事大学电气工程系; 滑铁卢大学系统设计工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究超长场景文本识别问题,通过双轴诊断发现编码器宽度是主要失败模式,提出无需训练的旋转缩放等修复方法,推理时通过特定裁剪和对齐方式弥合差距,在长文本基准上取得高准确率,发布相关工具和代码。
AI 中文摘要
场景文本识别(STR)模型几乎完全基于最多25个字符的单词裁剪进行训练,然而实际部署(如标识、产品标签、密集字幕)需要读取更长的文本。本文对这种失败情况进行了诊断并加以解决。诊断将超长失败分为两个同时外推的轴(编码器的宽度轴和解码器的时间轴),并表明编码器宽度而非解码器长度是主要失败模式。表示侧修复只能带来部分缓解:无需训练的旋转缩放最多可恢复2 - 4个字符错误率(CER)点,加权微调方法可恢复6 - 8个点并提高标准基准准确率,但长文本基准(LTB)上的单词准确率仍接近零,因为剩余差距在于解码机制而非表示。然后在推理时,在未修改的单词级检查点上弥合该差距:将长图像在模型训练宽度处切成重叠裁剪,每个独立解码且符合分布,通过几何锚定编辑距离对齐拼接读取结果。此过程在两个基础检查点上在LTB上达到42.79 - 43.05%的桶平均单词准确率,与已发表的最优水平(41.57%)匹配,并在最难的桶上高出11 - 12个点,与普通解码耗时相同;应用于公共PARSeq检查点时达到47.11%。一旦应用分块,微调不再有帮助:仅解码侧修复就可与专门构建的架构相匹配。我们发布了诊断工具和实现代码。
英文摘要
Scene Text Recognition (STR) models are trained almost exclusively on word crops of at most 25 characters, yet real deployments (signage, product labels, dense captions) require reading much longer text. This paper diagnoses that failure and then closes it. The diagnosis separates out-of-length failure into two simultaneously extrapolating axes (the encoder's width axis and the decoder's time axis) and shows that encoder width, not decoder length, is the dominant failure mode. Representation-side fixes bring only partial relief: training-free rotary rescalings recover at most 2-4 points of character error rate (CER), and a weighted fine-tuning recipe recovers 6-8 points while improving standard-benchmark accuracy, yet word accuracy on the Long Text Benchmark (LTB) stays near zero, because the residual gap lies in the decoding mechanism rather than the representation. We then close that gap at inference time, on an unmodified word-level checkpoint: the long image is sliced into overlapping crops at the model's training width, each decoded independently and in-distribution, and the reads stitched by geometry-anchored edit-distance alignment. This procedure reaches 42.79-43.05% bucket-average word accuracy on LTB across two base checkpoints, matching the published state of the art (41.57%) and beating it by 11-12 points on the hardest bucket, at wall-clock parity with plain decoding; applied unchanged to the public PARSeq checkpoint it reaches 47.11%. Once chunking is applied fine-tuning no longer helps: the decoding-side fix alone matches purpose-built architectures. We release the diagnosis harness and implementation.
Comments39 pages, 10 tables, includes supplementary material (folded in as Appendix S1-S4). Code: https://github.com/zobeirraisi/chunk-decode-stitch