评估天城文OCR后校正的上下文学习与检索策略
Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction
浏览论文内容
中文总结 AI 辅助
本研究首次系统评估LLM在天城文OCR后校正中的上下文学习,提出CharBM25检索策略,在印地语和马拉地语上显著优于随机选择,且与12B+模型结合实现免训练可靠校正。
中文摘要 AI 辅助
使用大型语言模型(LLMs)的上下文学习为免训练的后OCR校正提供了一条引人注目的路径,但其对天城文的有效性尚未被探索。我们首次对LLMs(3B-32B)在印地语和马拉地语的后OCR校正中进行了系统评估,比较了三种上下文示例检索策略:领域随机选择、密集语义检索以及我们提出的CharBM25,后者通过OCR输入上的字符n-gram BM25相似性检索示例,以针对测试句子的共享错误模式。在跨越五个新闻领域的20,000句基准测试中,检索策略是校正质量的决定性因素:CharBM25在印地语上比领域随机选择高出2.8-4.0个百分点绝对WER,在马拉地语上高出2.9-3.8个百分点,使用字符三元组,其始终优于二元组和一元组。规模主导性能:Gemma-3-27B在CharBM25-5下实现了印地语55.0%和马拉地语33.3%的WER降低。少样本增益受容量限制:低于8B的模型不能可靠地优于OCR基线,而在马拉地语上,最小的模型(3B)使更多句子退化而非改进。在所有规模上,马拉地语比印地语更难校正,反映了其更大的形态复杂性。这些发现确立了CharBM25作为一种有效的、无需GPU的检索策略,在可忽略的计算成本下匹配或超过密集检索,并表明将其与12B+参数的通用的LLM结合,可以提供可靠的、免训练的天城文后OCR校正,而无需任务特定的微调。数据集:此https URL
英文摘要
In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, comparing three in-context example retrieval strategies: domain-random selection, dense semantic retrieval, and our proposed CharBM25, which retrieves examples by character n-gram BM25 similarity over OCR inputs to target shared error patterns with the test sentence. Across a 20,000-sentence benchmark spanning five news domains, retrieval strategy is the decisive factor in correction quality: CharBM25 outperforms domain-random selection by 2.8-4.0pp absolute WER on Hindi and 2.9-3.8pp on Marathi, using character trigrams, which consistently outperform bigrams and unigrams. Scale dominates performance: Gemma-3-27B achieves WER reductions of 55.0% for Hindi and 33.3% for Marathi under CharBM25-5. Few-shot gains are capacity-gated: models below 8B do not reliably improve over the OCR baseline, and on Marathi the smallest models (3B) degrade more sentences than they improve. Marathi is persistently harder to correct than Hindi across all scales, reflecting its greater morphological complexity. These findings establish CharBM25 as an effective, GPU-free retrieval strategy that matches or exceeds dense retrieval at negligible computational cost, and show that combining it with a general-purpose LLM of 12B+ parameters delivers reliable, training-free Devanagari post-OCR correction without task-specific fine-tuning. Dataset: https://huggingface.co/datasets/AbhishekBhandari/Devanagari-OCR-ICL-Benchmark
发表机构
- Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)
机构由 AI 辅助整理,请以论文原文为准。