arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33603cs.CVcs.AI

ViCoR:通过空间对齐验证与可执行修订实现可靠的分子结构提取

ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision

  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yujian Yuan, Xin Cai, Yufan Chen, Jiaxin Xu, Mengdi Liu, Zhichao Tan, Long Chen, Hanyu Gao

AI总结:

ViCoR提出先修复后拒绝的框架,通过空间对齐验证与可执行修订提升光学化学结构识别的可靠性,显著提高准确率并改善下游化学任务性能。

AI中文摘要:

可靠的光学化学结构识别(OCSR)对于从科学文献中构建高质量化学数据至关重要,然而即使是微小的识别错误也可能传播到化学数据库和下游模型中。在实践中,识别出的结构在使用前通常需要人工检查和修正,这使得大规模数据整理成本高昂且难以扩展。因此,我们研究选择性结构识别(SSR),这是一种识别后处理设置,能够自动生成可靠的结构化输出,同时拒绝未解决的情况。仅选择的方法可以通过拒绝来提高可靠性,但无法产生超出基础识别器输出的额外正确结果。我们提出ViCoR,一种先修复后拒绝的迭代验证与修订框架。其关键思想是使观察-预测对应关系显式化:坐标保持渲染建立了源图像与预测结构之间的空间对应关系,而索引锚定将局部视觉差异映射为可执行的图编辑,无需重新生成完整结构。一个共享的视觉语言模型(VLM)从验证到修订逐步训练。在两个真实世界的OCSR基准上,ViCoR将整体准确率从73.53%提升至88.26%,从61.83%提升至84.32%,同时在85%至89%的覆盖率下实现了超过97%的接受准确率。所得的分子数据进一步将反应提取的F1分数提高了15.5个百分点,并将文献来源的反应预测准确率提高了7.7和5.8个百分点,证明了自动化可靠性控制对科学数据整理和下游化学学习的价值。

英文摘要:

Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection and correction before use, making large-scale data curation costly and difficult to scale. We therefore study Selective Structure Recognition (SSR), a post-recognition setting that automatically produces reliable structured outputs while rejecting unresolved cases. Selection-only approaches can improve reliability by rejection, but cannot create additional correct outputs beyond those produced by the base recognizer. We propose ViCoR, a repair-before-rejection framework for iterative VerIfiCatiOn and Revision. Its key idea is to make observation-prediction correspondence explicit: coordinate-preserving rendering establishes spatial correspondence between the source image and predicted structure, while index anchoring maps localized visual discrepancies to executable graph edits without full-structure regeneration. A shared VLM is progressively trained from verification to revision. On two real-world OCSR benchmarks, ViCoR improves overall accuracy from 73.53\% to 88.26\% and from 61.83\% to 84.32\%, while achieving over 97\% accepted accuracy at 85--89\% coverage. The resulting molecular data further improve reaction-extraction F1 by 15.5 points and literature-sourced reaction prediction accuracy by 7.7 and 5.8 points, demonstrating the value of automated reliability control for scientific data curation and downstream chemical learning.

↑