arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

嵌入旋转不变性的可证明多方向场景文本识别

Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition

Zhibin Ma, Pengwen Dai, Yi Liu, Xugong Qin, Chenyun Yu, Xiaochun Cao

arXiv 2608.10684首次发表:更新:

发表机构

Sun Yat-sen University; Baidu Inc.; Nanjing University of Science and Technology(中山大学; 百度公司; 南京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出带理论保证的端到端旋转不变场景文本识别网络RISTER,通过编码器嵌入旋转等变性、解码器嵌入旋转不变性,在多方向场景文本识别任务上实现最优性能且无额外推理成本。

AI 中文摘要

多方向文本在真实场景中普遍存在,仍是场景文本识别(STR)的主要挑战。现有感知旋转的方法会显式估计文本方向,但因缺乏理论保证,易出现误差累积、计算成本增加且强依赖数据。本研究将旋转不变性融入STR框架以解决这些局限。具体而言,采用编码器-解码器架构,在编码器中嵌入旋转等变性、解码器中嵌入旋转不变性,构建完全旋转不变的网络。解码器侧,首次识别并证明交叉注意力机制的旋转不变性,以此构建旋转不变文本解码器,能以旋转不变方式将视觉特征映射为输出文本;编码器侧,提出旋转等变的局部-全局提取网络,将深度等变卷积与自注意力结合,实现旋转等变特征提取,同时建模字符间依赖关系并保留细粒度视觉细节。整合编码器与解码器后,得到端到端的旋转不变场景文本识别网络(RISTER)。RISTER具备带理论保证的旋转不变性,可增强多方向样本的鲁棒性,且不会引入额外推理计算或依赖数据驱动的方向校正。实验表明,RISTER在标准及多方向基准上均达到最优性能,在通用多方向数据集上的准确率较次优模型高出4.0个百分点。

英文摘要

Multi-oriented text is ubiquitous in real-world scenes and remains a major challenge for scene text recognition (STR). Existing rotation-aware methods explicitly estimate text orientation. However, due to the lack of theoretical guarantees, they are prone to error accumulation, increased computational cost, and strong reliance on data. In this work, we incorporate rotation invariance into the STR framework to address these limitations. Specifically, we adopt an encoder-decoder architecture, embedding rotation equivariance in the encoder and rotation invariance in the decoder to construct a fully rotation-invariant network. On the decoder side, we first identify and prove the rotation-invariant property of the cross-attention mechanism and use it to formulate a rotation-invariant text decoder that maps visual features to output text in a rotation-invariant manner. On the encoder side, we propose a rotation-equivariant local-global extraction network that integrates deep equivariant convolutions with self-attention, enabling rotation-equivariant feature extraction while modeling inter-character dependencies and preserving fine-grained visual details. By integrating the encoder and decoder, we obtain an end-to-end Rotation-Invariant Scene Text Recognition network (RISTER). RISTER provides rotation invariance with theoretical guarantees, enhancing robustness on multi-oriented samples without introducing additional inference computation or relying on data-driven orientation correction. Experiments show that RISTER achieves state-of-the-art performance on both standard and multi-oriented benchmarks, surpassing the second-best model by 4.0 percent in accuracy on the general multi-oriented dataset.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑