arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

本体引导、推理器验证的科学人工智能大语言模型推理评估基准

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler

arXiv 2610.00682首次发表:更新:

发表机构

Siemens AG; Technical University of Munich(西门子股份公司; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种从OWL 2本体自动生成多项选择题基准的流水线,通过推理器验证干扰项错误性,在三个本体上生成大量题目,评估六个LLM准确率41.1%-76.8%,为科学AI逻辑推理提供可靠诊断基准。

AI 中文摘要

大语言模型(LLM)日益支撑着科学人工智能应用中基于结构化知识的推理,从生物医学问答到材料信息学。然而,其逻辑推理常常存在不足,产生在这些场景中不可接受的事实性错误。可靠的评估仍具挑战性:人工数据集构建扩展性差,而基于LLM的生成则可能嵌入其旨在衡量的缺陷。高质量的基准必须将正确和错误的标注示例都植根于明确的背景知识中,并可由标准推理器进行形式化验证。我们提出了一种流水线,可从任何充分公理化的OWL 2本体自动生成基于本体的多项选择题(MCQ)基准,正确答案在设计上植根于本体。干扰项通过扰动类定义公理的右侧类表达式生成,其错误性由OWL推理器通过蕴含检查进行形式化验证。我们在三个本体上评估该流水线:Pizza(小型、学术)、PMDco(复杂、材料科学)和DOID(大型、生物医学),分别生成112、2,491和15,216个MCQ。干扰项涵盖从类不可满足性到弱化子类关系的四个语义类别,能够对特定推理失败进行诊断性评估。项目达到自然语言质量标准:平均LLM评判得分为5分制中的4.02、4.36和3.36,确认了流畅性;正确答案与干扰项的相似度高于0.8,表明错误选项不能仅凭表面形式被排除。六个LLM在零样本设置下评估,准确率为41.1%-76.8%,远高于25%的随机猜测基线,表明这些基准具有挑战性和区分度。这项工作朝着更可靠地评估科学人工智能逻辑推理的基准迈出了一步。

英文摘要

Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.

Comments17 pages, 2 figures. Accepted at the AI Data Readiness for Scientific Discovery (AIDaR) Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026), Paris

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑