arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BanglaWild:面向OCR和视觉-语言模型的野外孟加拉语场景文本识别基准

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

Sadab Shiper, Tawsif Tashwar Dipto, Mir Md Inzamam, Eshat Tanzeem

arXiv 2608.03884首次发表:更新:

AI 中文总结

研究发现现有孟加拉语场景文本识别基准的不足,推出含2535张图像的BanglaWild基准,评估15个VLMs和3个OCR系统等,揭示视觉误识别为主要错误来源等结论,代码数据将公开。

AI 中文摘要

野外孟加拉语场景文本识别在很大程度上未得到充分测量:现有资源针对手写文档或受限的招牌解析,仅报告聚合编辑距离指标,且仅评估传统OCR或视觉-语言模型(VLMs),从未在同一野外数据上同时评估两者。为解决这一差距,我们推出BANGLAWILD,这是一个包含2535张孟加拉语场景文本图像的基准,每张图像都配有逐字黄金转录本、两个分类轴、四个诊断属性,以及当图像内文本偏离标准拼写时的正字法标准形式。我们在三种提示策略下评估15个VLMs和3个传统OCR系统,用LoRA微调6个开源模型,并以LLM作为评判者的评估补充编辑距离指标。我们的结果显示,同一系列中更大的模型并未优于更小的模型,这一差距持续存在;我们的15类错误分类显示,在最强系统中,视觉误识别占错误的约60%,而与连词相关的错误贡献不足2%,这挑战了孟加拉语OCR研究中的长期假设,且相同的视觉主导特征在各架构中均存在,包括可靠读取孟加拉语的传统基线。提示语言主要影响跨脚本漂移,LoRA减少了弱模型的灾难性故障,但未提升已表现良好模型的上限。代码和数据将公开发布。

英文摘要

In-the-wild Bengali scene text recognition is largely unmeasured: existing resources target handwritten documents or constrained sign-board parsing, report only aggregate edit-distance metrics, and evaluate either conventional OCR or VLMs, never both on the same in-the-wild data. To address this gap, we introduce BANGLAWILD, a benchmark of 2,535 Bengali scene text images, each paired with a verbatim gold transcription, two categorical axes, four diagnostic attributes, and an orthographically standard form where the in-image text deviates from canonical spelling. We evaluate fifteen VLMs and three conventional OCR systems under three prompting strategies, fine-tune 6 open-source models with LoRA, and complement edit-distance metrics with an LLM-as-a-Judge evaluation. Our results reveal a persistent gap in which larger models within the same family do not outperform smaller ones. Our fifteen-class error taxonomy shows that visual mis-recognition accounts for ~60% of errors in the strongest systems, while conjunct-related errors contribute under 2%, challenging a long-standing assumption in Bengali OCR research; the same visual dominant profile also holds across architectures, including the one conventional baseline that reads Bengali reliably. Prompt language mainly affects cross-script drift and LoRA reduces catastrophic failures in weak models without lifting the ceiling on already competent ones. Code and data will be publicly released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑