arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MolParser-Mobile:面向大规模化学文献挖掘的超快OCSR系统

MolParser-Mobile: Ultrafast OCSR System for Large-Scale Chemical Literature Mining

Xi Fang, Haocheng Lu, Han Lyu, Chengxiang Luo, Linfeng Zhang, Guolin Ke

arXiv 2609.05807首次发表:更新:

发表机构

DP Technology(DP Technology)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MolParser-Mobile,一个AutoML优化的轻量级端到端OCSR框架,仅含9.98M参数,在单块RTX 4090D GPU上实现每秒1520个分子的吞吐量,同时保持甚至超越现有方法的识别精度,解决了大规模化学文献挖掘中的推理瓶颈问题。

AI 中文摘要

光学化学结构识别(OCSR)是化学文献挖掘的基础组成部分,能够支持分子数据库构建、反应提取以及人工智能驱动的科学发现。尽管近年来基于深度学习的方法在识别精度上取得了显著进展,但推理吞吐量仍然是限制其网络规模部署的关键瓶颈。为解决这一挑战,我们提出了MolParser-Mobile,一个经AutoML优化的轻量级端到端OCSR框架。MolParser-Mobile仅包含998万(9.98M)个参数,在单个NVIDIA RTX 4090D GPU上达到了每秒1520个分子的吞吐量。尽管设计紧凑,它在多个基准测试上保持了具有竞争力的识别精度,并在若干基准上取得了更优的性能。

英文摘要

Optical Chemical Structure Recognition (OCSR) is a fundamental component of chemical literature mining, enabling molecular database construction, reaction extraction, and AI-driven scientific discovery. Despite substantial progress in recognition accuracy with recent deep learning-based methods, inference throughput remains a critical bottleneck that limits web-scale deployment. To address this challenge, we propose MolParser-Mobile, an AutoML-optimized lightweight end-to-end OCSR framework. MolParser-Mobile contains only 9.98M parameters, while reaching a throughput of 1,520 molecules per second on a single NVIDIA RTX 4090D GPU. Despite its compact design, it maintains competitive and, on several benchmarks, superior recognition accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑