发表机构
Chabahar Maritime University(恰巴哈尔海事大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对场景文本识别中基于Transformer方法依赖一维位置编码忽略文本二维空间结构的问题,提出2D - RoPE - STR,通过各向异性维度分配和旋转耦合扩展到交叉注意力,提升不规则布局识别效果,且无参数并无需架构重设。
AI 中文摘要
场景文本识别(STR)因文本外观的多样性(包括曲率、旋转和透视失真)而仍然具有挑战性。最近基于Transformer的方法表现良好,但通常依赖于忽略文本图像二维空间结构的一维位置编码。视觉Transformer存在旋转位置嵌入(RoPE)的轴向二维扩展,但它们假设图像内容大致为方形且各向同性,并且仅在编码器自注意力内应用旋转。场景文本违反了这两个假设:裁剪明显各向异性,并且STR模型是编码器 - 解码器结构,因此解码器必须通过交叉注意力将其查询与编码器的二维布局相关联。我们引入了2D - RoPE - STR,它通过(1)与文本宽高比匹配的各向异性行/列维度分配,以及(2)将旋转耦合扩展到编码器 - 解码器交叉注意力,使自回归解码步骤能够根据其二维布局关注编码器令牌,这是先前仅编码器公式未解决的设置。这两个变化基本上都是无参数的,并且除了位置编码模块之外不需要架构重新设计。我们还引入了一种诊断协议(一个仅隔离位置编码的受控消融对、图像级网络 - 获胜者不一致分析和编码器注意力可视化),该协议确定相对二维位置在何处以及为何有帮助:弯曲、旋转和透视失真的布局,其中阅读顺序偏离直线水平线。在六个标准基准(IIIT5K、SVT、ICDAR 2013、ICDAR 2015、CUTE80、SVTP)上,增益集中在这些不规则布局上,消融将每个设计选择与一维RoPE以及二维正弦和可学习替代方案进行了对比。
英文摘要
Scene Text Recognition (STR) remains challenging due to the diversity of text appearances, including curvature, rotation, and perspective distortion. Recent Transformer-based approaches perform well but usually rely on one-dimensional positional encodings that ignore the 2D spatial structure of text images. Axial 2D extensions of Rotary Position Embedding (RoPE) exist for vision Transformers, but they assume roughly square, isotropic image content and apply the rotation only within encoder self-attention. Scene text violates both assumptions: crops are markedly anisotropic, and STR models are encoder-decoder, so the decoder must relate its queries to the encoder's 2D layout through cross-attention. We introduce 2D-RoPE-STR, which adapts axial 2D-RoPE to this setting through (1) an anisotropic row/column dimension allocation matched to the aspect ratio of text, and (2) an extension of the rotary coupling into encoder-decoder cross-attention, letting autoregressive decoding steps attend to encoder tokens by their 2D layout, a setting not addressed by prior encoder-only formulations. Both changes are essentially parameter-free and require no architectural redesign beyond the positional-encoding module. We further introduce a diagnostic protocol (a controlled ablation pair isolating only the positional encoding, an image-level net-win disagreement analysis, and encoder attention visualization) that identifies where and why relative 2D position helps: curved, rotated, and perspective-distorted layouts where reading order departs from a straight horizontal line. On six standard benchmarks (IIIT5K, SVT, ICDAR 2013, ICDAR 2015, CUTE80, SVTP), gains concentrate on exactly these irregular layouts, with ablations isolating each design choice against 1D RoPE and 2D sinusoidal and learnable alternatives.
Comments17 pages, 3 figures. Under review at the International Journal on Document Analysis and Recognition (IJDAR)