发表机构
Friedrich-Alexander-Universität Erlangen–Nürnberg(埃尔朗根-纽伦堡弗里德里希-亚历山大大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对遗产数字化需求,对比8款最高7B参数的开源轻量级VLMs在文档OCR转结构化JSON上的表现,为遗产机构提供可持续、私有且高性能的替代方案及操作指导。
AI 中文摘要
尽管大量闭源视觉-语言模型(VLMs)在文档理解领域设定了强大的基准,但由于数据自主性顾虑、持续成本以及超大规模计算的环境足迹,它们对商业API的依赖限制了其在机构档案中的应用。这在遗产数字化领域尤为突出,其中的文档包含历史手写体、特定领域术语(如珠宝、史前学、建筑学)以及需要高维结构化提取的非标准布局。我们针对三个大学遗产馆藏,开展了八项开源轻量级VLMs(参数规模最高达7B)用于光学字符识别(OCR)转结构的对比研究。给定一张文档图像,模型必须提取文本并生成符合 schema 的JSON,以支持自动验证和下游使用。我们在感知约束的协议下,针对零样本、少样本和微调设置评估模型,使用字符错误率(CER)、近似归一化莱文斯坦相似度(ANLS*)和平均精度均值F1(mAP-F1)衡量提取保真度和结构化输出质量。相较于微调基线,我们进一步测试了(i)超参数优化、(ii)经典图像预处理(照度扁平化、去噪和CLAHE)以及(iii)多阶段训练的独立影响。最后,我们分析了特定数据集微调与单一多数据集检查点之间的权衡,其中联合训练可使单个模型在多个馆藏间运行,但会在不同数据集间转移性能。总体而言,我们表明经过精心适配的、参数规模最高达7B的VLMs可提供可持续、私有且高性能的替代方案,替代手动转录或商业黑箱系统,同时为寻求机构可控OCR转JSON提取的遗产机构提供可操作的指导。
英文摘要
While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image, models must extract text and generate schema-compliant JSON, enabling automatic validation and downstream use. We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS*), and mean Average Precision F1 (mAP-F1). Against a fine-tuning baseline, we further test the independent impact of (i) hyperparameter optimization, (ii) classical image preprocessing (illumination flattening, denoising, and CLAHE), and (iii) multi-stage training. Finally, we analyze the trade-off between dataset-specific fine-tuning and a single multi-dataset checkpoint, where joint training enables one model to operate across collections but can shift performance between datasets. Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.
Comments17 pages. Published in Document Analysis and Recognition - ICDAR 2026, LNCS vol. 16974, Springer. Code: https://github.com/uddipan77/Analysis-of-Lightweight-Vision-Language-Models-for-Document-OCR-and-Structured-Output-Generation
Journal refDocument Analysis and Recognition - ICDAR 2026, Lecture Notes in Computer Science, vol. 16974, pp. 502-519, Springer, 2027
DOI:10.1007/978-3-032-36039-7_30