发表机构
SES AI Corporation(SES AI公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对光学化学结构识别中合成与真实数据的性能差距,通过微调不同视觉语言模型,发现标注真实训练数据可大幅提升性能,且模型与适配策略的选择需结合目标任务评估。
AI 中文摘要
数百万个化学结构仅以绘图形式出现在专利和论文中,大规模利用这些信息需要读取这些绘图。光学化学结构识别(OCSR)在合成图像上几乎已被解决,但在真实文档上仍然困难:初始识别模型Qwen2.5-VL-7B在合成渲染图上的准确率超过91%,但在三个真实基准测试(ACS、CLEF-IP、USPTO)上的准确率低于16%。为确定改进的主要来源,研究人员在合成渲染结构与来自专利、期刊图表和手绘集合的标注真实描绘的混合数据上,对21个识别模型进行微调,改变了视觉语言模型(VLM)的基础模型、真实训练数据的比例以及视觉塔适配策略。标注真实训练图像带来的改进最大。对于Qwen2.5-VL,ACS精确匹配率从无真实数据时的0.15,在真实数据占比9.5%时升至0.37,在真实数据占比50.2%时升至0.46;对三个基础模型的对照实验重现了这一趋势。相比之下,视觉塔LoRA对Qwen无作用(+0.00,配对p值=1.00),对InternVL3-8B有显著帮助(+22.8至+34.6个百分点),对GLM-4.1V-9B有适度帮助(+1.0至+9.6个百分点),因此其价值取决于基础模型。最佳配置在干净渲染图上达到0.96的精确匹配率,在ACS、CLEF-IP、UOB和USPTO上分别达到0.49、0.65、0.84和0.76的精确匹配率。基础模型之间的差距在无真实数据时最大(0.21),在真实数据占比70%时缩小至0.06,且排名发生变化;因此必须同时选择基础模型和真实数据混合比例。对手写图像到LaTeX识别和图表到表格转换的小规模实验表明,基础模型的排名在化学领域之外也会变化。更普遍地,视觉结构识别的模型和适配选择应在目标任务上进行评估。
英文摘要
Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR appears nearly solved on synthetic images yet remains difficult on real documents: the starting recognizer, Qwen2.5-VL-7B, exceeds 91% accuracy on synthetic renders but falls below 16% on three real-world benchmarks (ACS, CLEF-IP, USPTO). To identify the main source of improvement, 21 recognizers were fine-tuned on mixtures of synthetically rendered structures and labeled real depictions from patents, journal figures, and hand-drawn collections, varying the vision language model (VLM) base, the fraction of real training data, and the vision-tower adaptation strategy. Labeled real training images make the largest difference. For Qwen2.5-VL, ACS exact match rises from 0.15 with no real data to 0.37 at 9.5% and 0.46 at 50.2%; a controlled experiment across three base models reproduces the trend. A vision-tower LoRA, in contrast, does nothing for Qwen (+0.00, paired p=1.00), substantially helps InternVL3-8B (+22.8 to +34.6 pt), and modestly helps GLM-4.1V-9B (+1.0 to +9.6 pt), so its value depends on the base model. The best configuration reaches 0.96 exact match on clean renders and 0.49, 0.65, 0.84, and 0.76 on ACS, CLEF-IP, UOB, and USPTO, respectively. Gaps between base models are largest without real data (0.21), shrink to 0.06 at 70% real data, and reorder the ranking; base model and real-data mixture must therefore be selected together. Small-scale experiments on handwritten image-to-LaTeX recognition and chart-to-table conversion show that base-model rankings also vary beyond chemistry. More generally, model and adaptation choices for visual structure recognition should be evaluated on the target task.