arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28857cs.CV

MEVL-STP:用于任意形状场景文本识别的多编码器与视觉语言模型

MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting

  • Rajiv Gandhi Institute of Petroleum Technology(拉吉夫·甘地石油技术学院)
  • Indian Institute of Technology (ISM) Dhanbad(印度理工学院(ISM)丹巴德分校)
  • University of Salford(索尔福德大学)

机构由 AI 辅助整理,请以论文原文为准。

Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote

AI总结:

提出两阶段场景文本识别方法,用多冻结视觉编码器融合分割生成多边形掩码,再经LoRA微调的视觉语言模型识别,在CTW1500上达到SOTA,无需合成数据。

AI中文摘要:

场景文本识别对于自然图像中的任意形状文本实例(如弯曲标志和密集的多方向字符)仍然具有挑战性,在这些情况下,紧密耦合的架构会将定位误差直接传播为识别失败。我们提出了一种两阶段流程,将多编码器分割与视觉语言模型识别相结合来解决此问题。在检测阶段,六个冻结的视觉编码器(CLIP、DINOv2、SigLIP、EVA-CLIP、SAM 和 ConvNeXt)提取涵盖语义、空间和纹理谱的互补特征,这些特征通过一个可训练的分层特征金字塔网络(带有通道注意力)进行融合,并通过深度监督的渐进式尺度扩展网络进行解码,以生成精确的实例级文本掩码。通过保持编码器冻结,它们独立学习的特征空间在融合期间保持正交,从而防止了在单骨干检测器中降低边界精度的特征同质化。检测阶段生成紧密的多边形掩码,这些掩码符合弯曲和任意方向文本的实际形状,而不是不可避免地包含背景内容的轴对齐矩形。在识别阶段,这些多边形掩码裁剪将目标文本与周围杂波隔离,使得通过低秩适配在多边形裁剪的场景文本上微调的 Qwen3-VL-8B-Instruct 模型能够专注于纯粹阅读文本,而不受相邻单词或背景噪声的干扰。在没有使用任何合成预训练数据的情况下,我们的方法在 CTW1500 上实现了 91.99% 的检测 F-measure 和 85.86% 的端到端 H-mean,创下了新的最先进水平,并在 Total-Text 和 ICDAR 2015 上取得了强劲性能,且无需任何合成训练数据。代码可在 https 此 URL 获取。

英文摘要:

Scene text spotting remains challenging for arbitrarily shaped text instances such as curved signs and dense multi-oriented characters in natural images, where tightly coupled architectures propagate localization errors directly into recognition failures. We present a two-stage pipeline that combines multi-encoder segmentation with vision-language model recognition to address this problem. In the detection stage, six frozen vision encoders (CLIP, DINOv2, SigLIP, EVA-CLIP, SAM, and ConvNeXt) extract complementary features spanning semantic, spatial, and texture spectra, which are fused through a trainable hierarchical Feature Pyramid Network with channel attention and decoded via a deep-supervision Progressive Scale Expansion network to generate precise instance-level text masks. By keeping the encoders frozen, their independently learned feature spaces remain orthogonal during fusion, preventing the feature homogenization that degrades boundary precision in single-backbone detectors. The detection stage produces tight polygon masks that conform to the actual shape of curved and arbitrarily oriented text, rather than axis-aligned rectangles that inevitably include background content. In the recognition stage, these polygon-masked crops isolate the target text from surrounding clutter, allowing a Qwen3-VL-8B-Instruct model, fine-tuned via Low-Rank Adaptation on polygon-cropped scene text, to focus purely on reading the text without interference from neighbouring words or background noise. Without any synthetic pretraining data, our method achieves 91.99% detection F-measure and 85.86% end-to-end H-mean on CTW1500, setting a new state of the art and achieving strong performance on Total-Text and ICDAR 2015 without any synthetic training data. Code is available at https://github.com/doubleblind-afk/MEVL-STP

补充信息

↑