arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Spike-HTR:用于手写文本识别的脉冲神经Transformer

Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition

Xiubo Liang, Jinxing Han, Yuke Li, Haoqi Zhu, Yu Zhao, Hongzhi Wang

arXiv 2608.01646首次发表:更新:

AI 中文总结

针对手写文本识别的计算不平衡问题,提出混合脉冲识别器Spike-HTR,通过InkCoder和CTC引导的长度缩减器优化,在多个数据集上取得低字符错误率。

AI 中文摘要

手写文本识别(HTR)存在两方面计算不平衡问题:图像中大部分像素为背景,宽度轴序列的许多位置以空白为主。这与脉冲神经网络(SNN)不匹配:手写被视为静态图像,而脉冲计算随时间步展开。我们提出Spike-HTR,一种混合脉冲识别器,可控制脉冲步数和深度序列混合器处理的宽度位置数量。为使静态图像适配短程脉冲推理,InkCoder将其转换为由粗到细的输入流,早期步骤覆盖广泛笔画区域,后期步骤强调更精细的笔画细节。为减少序列计算,受连接时序分类(CTC)引导的长度缩减器保留可能的字符或不确定位置,在深度混合前压缩长的空白主导片段。当T=2时,Spike-HTR仅在目标数据上训练,解码时不使用语言模型或词典,在IAM、LAM和READ2016数据集上分别达到验证集/测试集字符错误率(CER)为3.5/5.4、2.3/2.5和4.2/3.9。代码可在此https URL获取。

英文摘要

Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background, and many width-axis sequence positions are blank-dominated. This creates a mismatch for Spiking Neural Networks (SNNs): handwriting is observed as a static image, whereas spiking computation unfolds over timesteps. We propose Spike-HTR, a hybrid spiking recognizer that controls both the number of spiking steps and the number of width positions processed by the deep sequence mixer. To make a static image suitable for short-horizon spiking inference, InkCoder converts it into a coarse-to-fine input stream, where early steps cover broad stroke regions and later steps emphasize sharper stroke details. To reduce sequence computation, a CTC-guided length reducer keeps likely character or uncertain positions and compresses long blank-dominated stretches before deep mixing. With $T{=}2$, Spike-HTR trains only on target data, decodes without language models or lexicons, and reaches validation/test CERs of 3.5/5.4, 2.3/2.5, and 4.2/3.9 on IAM, LAM, and READ2016. Codes are available at https://github.com/QomolangmaH/SpikeHTR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑