发表机构
Pontifical Catholic University of Paraná(巴拉那州天主教大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对艺术字体识别的挑战,以WordArt-V1.5为基准,提出结合SVTRv2等模型的置信度感知集成与长单词优化方法,提升了识别准确率。
AI 中文摘要
艺术字体识别(Artistic Text Recognition, ATR)仍具挑战性,因为单词图像常包含装饰性字体、弯曲布局、类物体字符、杂乱元素及严重扭曲。本文以WordArt-V1.5为该场景的标准化基准,在统一协议下评估近期场景和艺术字体识别器。我们提出一种置信度感知集成方法,在官方训练集上微调SVTRv2、PARSeq和MAERec后将三者结合;该集成通过分歧位置的最小置信度选择预测结果,重点区分竞争假设中的字符。对于长单词(单个字符错误即可导致整体预测失效),我们添加基于Needleman-Wunsch对齐和词典引导修正的针对性优化阶段。在WordArt-V1.5 Test B划分上,所提系统达到89.90%的单词识别准确率,较最佳的单个微调模型提升1.77个百分点;长单词优化虽仅带来小幅全局增益,但对目标长单词子集的准确率提升2.72个百分点。最后,对所有剩余错误的分析显示,48.8%的错误与标注问题、视觉模糊或难以辨认的样本相关,凸显了诊断报告对未来ATR基准和模型的价值。我们的源代码可在该https链接获取。
英文摘要
Artistic Text Recognition (ATR) remains challenging because word images often combine decorative fonts, curved layouts, object-like characters, clutter, and severe distortions. This paper studies WordArt-V1.5 as a standardized benchmark for this setting and evaluates recent scene and artistic text recognizers under a common protocol. We propose a confidence-aware ensemble that combines SVTRv2, PARSeq, and MAERec after fine-tuning on the official training split. The ensemble selects predictions using the minimum confidence over disagreement positions, emphasizing characters that separate competing hypotheses. For long words, where a single character error can invalidate the whole prediction, we add a targeted refinement stage based on Needleman-Wunsch alignment and lexicon-guided correction. On the WordArt-V1.5 Test B split, the proposed system reaches 89.90% Word Recognition Accuracy, improving the best individual fine-tuned model by 1.77 percentage points. The long-word refinement produces a modest global gain, but improves the targeted long-word subset by 2.72 percentage points. Finally, an error analysis of all remaining mistakes shows that 48.8% are associated with labeling issues, visual ambiguity, or illegible samples, highlighting the value of diagnostic reporting for future ATR benchmarks and models. Our source code is available at https://github.com/lucas-azdias/Artistic-Text-Recognition/.
CommentsAccepted for presentation at the 2026 Conference on Graphics, Patterns and Images (SIBGRAPI)