AI 中文总结
针对手写文本识别的计算不平衡问题,提出混合脉冲识别器Spike-HTR,通过InkCoder和CTC引导的长度缩减器优化,在多个数据集上取得低字符错误率。
AI 中文摘要
手写文本识别(HTR)存在两方面计算不平衡问题:图像中大部分像素为背景,宽度轴序列的许多位置以空白为主。这与脉冲神经网络(SNN)不匹配:手写被视为静态图像,而脉冲计算随时间步展开。我们提出Spike-HTR,一种混合脉冲识别器,可控制脉冲步数和深度序列混合器处理的宽度位置数量。为使静态图像适配短程脉冲推理,InkCoder将其转换为由粗到细的输入流,早期步骤覆盖广泛笔画区域,后期步骤强调更精细的笔画细节。为减少序列计算,受连接时序分类(CTC)引导的长度缩减器保留可能的字符或不确定位置,在深度混合前压缩长的空白主导片段。当T=2时,Spike-HTR仅在目标数据上训练,解码时不使用语言模型或词典,在IAM、LAM和READ2016数据集上分别达到验证集/测试集字符错误率(CER)为3.5/5.4、2.3/2.5和4.2/3.9。代码可在此https URL获取。
英文摘要
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background, and many width-axis sequence positions are blank-dominated. This creates a mismatch for Spiking Neural Networks (SNNs): handwriting is observed as a static image, whereas spiking computation unfolds over timesteps. We propose Spike-HTR, a hybrid spiking recognizer that controls both the number of spiking steps and the number of width positions processed by the deep sequence mixer. To make a static image suitable for short-horizon spiking inference, InkCoder converts it into a coarse-to-fine input stream, where early steps cover broad stroke regions and later steps emphasize sharper stroke details. To reduce sequence computation, a CTC-guided length reducer keeps likely character or uncertain positions and compresses long blank-dominated stretches before deep mixing. With $T{=}2$, Spike-HTR trains only on target data, decodes without language models or lexicons, and reaches validation/test CERs of 3.5/5.4, 2.3/2.5, and 4.2/3.9 on IAM, LAM, and READ2016. Codes are available at https://github.com/QomolangmaH/SpikeHTR.