arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

R-GroundBench:Markush分子编辑中R基团定位的诊断基准

R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

Xin Wang, Zichuan Ying, Xinna Lin, Junqi Zhang, Hanyi Xiong, Tianyu Gao, Hairong Zhang, Qixiang Hua, Botian Shi, Zhenhailong Wang, Kaicheng Yu

arXiv 2610.00700首次发表:更新:

发表机构

Westlake University; The University of Hong Kong; Shanghai Innovation Institute; Zhejiang University; Sichuan University; Shanghai Artificial Intelligence Laboratory; The Hong Kong University of Science and Technology (Guangzhou); University of Illinois Urbana-Champaign(西湖大学; 香港大学; 上海创新研究院; 浙江大学; 四川大学; 上海人工智能实验室; 香港科技大学(广州); 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Markush分子编辑中R基团定位缺失的问题,构建了基于真实专利的诊断基准R-GroundBench,通过VQA与生成任务揭示现有模型在困难场景下性能大幅下降,缺乏可靠定位与执行能力。

AI 中文摘要

近期人工智能在科学发现领域的进展使得分子理解与设计成为可能,然而,对不完整化学表示进行推理仍然是一个挑战。Markush结构通过可变的R基团占位符(如R₁、R₂、X等)编码分子家族,在药物专利中普遍存在,需要在分子、文本和化学上下文中进行定位。然而,现有的分子-语言基准主要关注完全指定的分子,R基团定位在很大程度上未被探索。我们引入了R-GroundBench,这是一个基于真实专利Markush结构构建的诊断基准,包含一个具有受控难度和模态划分的多选题(VQA)轨道,以及一个开放式的生成轨道。实验结果揭示了识别与分子执行之间的显著差距:虽然模型在简单VQA上达到超过90%的准确率,但在去除捷径的困难VQA上性能降至56%至66%。跨领域视觉语言模型(VLMs)仍然不可靠,尽管有领域特定的提示,在困难VQA上仅达到25.7%至46.2%的准确率。此外,生成轨道的精确匹配率在大多数模型上低于20%,当视觉输入被移除时低于8%。这些发现表明,当前的AI系统缺乏对Markush编辑的可靠定位与执行能力,凸显了AI驱动科学发现所面临的挑战。

英文摘要

Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely unevaluated.We introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation track.Our results reveal a substantial gap between recognition andmolecular grounding.While models achieve over 90\% accuracy on Easy VQA, performance drops to56--66\% on Hard VQA when shortcuts are controlled.Chemical-domain VLMs also remain unreliable, achieving only 25.7--46.2\% on HardVQA despite domain-specific pretraining.Moreover, Generation Exact Match remains below 20\% for most models and below8\% when visual input is required.These findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑