发表机构
Huazhong University of Science and Technology; Adobe Research(华中科技大学; 奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本spotting中识别与定位能力难以兼顾的问题,提出SupGRPO联合训练策略,结合SFT与GRPO并设计基于匹配的在线SFT,在艺术化文本数据集上同时提升识别与检测性能。
AI 中文摘要
文本spotting需要同时具备准确的文本识别和精确的空间定位能力。当前专门的spotter在自然场景中预测紧致边界框方面表现出色,但在复杂或艺术化文本上表现不佳,而多模态大语言模型(MLLMs)拥有强大的识别能力,但在定位方面仍然薄弱。为了使文本spotter具备通用且强大的识别能力,并最大化其定位能力,我们探索了两种基于MLLM的微调方法:监督微调(SFT)和基于组相对策略优化(GRPO)的强化学习微调。一个有趣的发现是,SFT在增强识别方面不如GRPO有效,而GRPO在增强检测方面不如SFT有效。为了互补各自的不足,我们引入了一种联合训练策略SupGRPO,该策略同时使用SFT和GRPO优化模型。SupGRPO采用专门设计的奖励函数,并开发了一种仅应用于坐标token的基于匹配的在线SFT。它既缓解了GRPO的奖励稀疏问题,又避免了SFT的实例顺序依赖问题。为了评估特别具有挑战性的案例,我们整理了ATS,一个用于艺术化文本spotting的数据集。实验表明,SupGRPO同时提升了文本识别和检测性能,并取得了优越的结果。我们的代码和数据集将在该https URL上发布。
英文摘要
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised spotters excel at predicting tight bounding boxes in natural scenes, but falter on complex or artistic text, whereas multimodal large language models (MLLMs) possess strong recognition capabilities yet remain weak at localisation. To equip the text spotter with general and powerful recognition capabilities and to maximize its localization ability, we explore two MLLM-based fine-tuning methods: Supervised Fine-Tuning (SFT) and reinforcement learning fine-tuning based on Group Relative Policy Optimisation (GRPO). An interesting finding is that SFT is less effective than GRPO at enhancing recognition, while GRPO is less effective than SFT at enhancing detection. To compensate for each other's shortcomings, we introduce a joint training strategy, SupGRPO, which simultaneously optimizes the model using both SFT and GRPO. SupGRPO employs the specially designed reward functions and develops a matching-based online SFT applied solely to coordinate tokens. It both mitigates the reward sparsity problem of GRPO and avoids the instance order dependency problem of SFT. To evaluate particularly challenging cases, we curate ATS, a dataset for artistic text spotting. Experiments demonstrate that SupGRPO improves both text recognition and detection, and attains superior performance. Our code and dataset will be released at https://github.com/Psycho-9/SupGRPO.
CommentsAccepted by ECCV 2026