发表机构
South China University of Technology; HiThink Research(华南理工大学; 海思思考研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多补丁场景文本定位范式的冗余噪声与定位歧义问题,提出以视觉为中心的单补丁文本定位框架 SPaTS,通过强化学习优化的单补丁选择等技术实现性能提升,显著优于前沿相关模型。
AI 中文摘要
场景文本定位需要文本识别与空间定位之间的高精度对齐。尽管视觉标记接地已成为多模态大语言模型(MLLM)的一种有前景的范式,但先前的多补丁范式常会引入冗余噪声和定位歧义,尤其针对密集或小型文本实例。为解决该问题,我们提出单补丁文本定位(SPaTS),这是一种以视觉为中心的框架,它将每个文本实例通过单个锚点视觉标记,再通过全图细化恢复几何信息。为在无需先验标签的情况下准确识别该锚点,我们引入单补丁选择性优化(SPaSO),这是一种使用补丁级奖励优化离散视觉标记选择的强化学习框架。为进一步提升表示鲁棒性和定位精度,我们引入方向嵌入对齐(DEA),通过解耦特征幅值与方向来抑制不稳定的幅值偏差,还引入补丁增强解码(PED),将路由的锚点与语言语义融合,并对全图特征图进行交叉注意力,以实现超越坐标空间替代物的几何感知边界回归。大量实验表明,SPaTS 始终显著优于前沿的闭源 MLLM 和 OCR MLLM,代码将很快发布。
英文摘要
Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code is available at https://github.com/eeNickTang/SPaTS.
Comments15 pages, 11 figures. Accepted to ACM Multimedia 2026
Journal refProceedings of the 34th ACM International Conference on Multimedia (MM '26), 2026