发表机构
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现场景文本识别在罕见词与罕见三元组组合上的准确率显著低于中心区域,定位故障源于自回归解码器的词汇先验,采用CTC解码的架构转变可有效缓解该问题,罕见输入长尾需架构变革而非增加容量。
AI 中文摘要
据报道,场景文本识别在6个标准基准上的准确率为89%-97%,该问题被广泛认为已接近饱和。我们提出了不同的视角:当针对参考语料库,按真实单词稀有度和字符n元组新颖性对相同测试图像进行联合分层时,在由此产生的5×5网格的“罕见词×罕见三元组”角落,9个英文专用识别器的准确率比q3/q3中心低10-18个百分点,且在我们测试的4种书写系统(拉丁文、汉字、汉字+假名、阿拉伯文)的13个(语言、模型)对中,全部13个对都存在“角落低于中心”的相同趋势。这种下降并非容量瓶颈导致:6倍视觉骨干网络规模扩展(CLIP4STR-Base 1.58亿参数→CLIP4STR-Huge 10亿参数,OpenCLIP ViT-H/14 LAION-2B)虽提升了所有基准的总准确率,但该压力角落的准确率未变(从86.9→86.5,在配对自助抽样噪声范围内)。4种收敛探测方法——分层探测、错误时的置信度、注意力重新平衡、跨脚本“提交vs弃权(不执行)”错误拆分——将故障定位到自回归解码器的词汇先验。随后我们探究现有技术能弥补多少差距:在16种非架构缓解措施中,平均q5/q5增益最大为+1.3个百分点,且无措施超过配对自助抽样噪声下限;唯一有效的干预措施是从自回归转向CTC解码的架构转变(SVTRv2,+2.5个百分点,p=0.02,n=474);置信度路由的AR-CTC集成带来方向一致的+0.6个百分点增益,仍在噪声范围内,其主导学习系数是各模型自身的最小softmax置信度,独立呼应上述机制。我们测试的所有配置均无法同时提升组合角落和总准确率,因此罕见输入长尾问题指向架构变革而非容量增加。
英文摘要
Scene text recognition is reported as 89--97% accurate on the six standard benchmarks, and the problem is widely treated as saturated. We present an alternative reading. When the same test images are stratified jointly by ground-truth word rarity and character n-gram novelty against a reference corpus, accuracy at the rare-word x rare-trigram corner of the resulting 5x5 grid drops 10--18 pt below the q3/q3 centre across nine English specialised recognisers, and the same direction (corner below centre) holds on all 13 of 13 (language, model) pairs we test across four writing systems (Latin, Han, Han+kana, Arabic). The drop is not a capacity bottleneck. A 6x vision-backbone scale-up (CLIP4STR-Base 158M -> CLIP4STR-Huge 1.0B, OpenCLIP ViT-H/14 LAION-2B) leads every benchmark in aggregate accuracy yet leaves the stress corner unchanged (86.9 -> 86.5, within paired-bootstrap noise). Four converging probes--layer-wise probing, confidence-when-wrong, attention re-balancing, and a cross-script commit-vs-abstain error split--localise the failure to the autoregressive decoder's lexical prior. We then ask how much of the gap existing techniques recover. Of 16 non-architectural mitigations, the largest mean q5/q5 gain is +1.3 pt and none clears the paired-bootstrap noise floor; the only intervention that does is the architectural shift from autoregressive to CTC decoding (SVTRv2, +2.5 pt, p=0.02, n=474). A confidence-routed AR-CTC ensemble adds a directionally consistent +0.6 pt that stays within noise, and its dominant learned coefficient is each model's own minimum-softmax confidence--independently echoing the mechanism above. No configuration we test improves both the compositional corner and aggregate accuracy. The rare-input long tail thus points to architectural change rather than added capacity.
CommentsPreprint. Under review. Main text plus appendix: 10 figures, 13 tables