arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以任务为中心的本体与确定性领域规则:AI辅助化学问题求解的可验证核心

A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving

Ibrokhimsho Abduchaborov

arXiv 2608.26164首次发表:更新:

AI 中文总结

本文提出ChemOntoRule,以任务为中心构建本体与确定性规则,在300道中学化学题上取得98.67%的匹配准确率,为AI辅助化学问题求解提供可验证的符号核心。

AI 中文摘要

大型语言模型能够解读自然语言化学问题,但其内部推理过程难以检查、约束与验证。本文提出ChemOntoRule,这是一个用于AI辅助中学化学问题求解的概念验证符号核心。核心设计选择是采用以任务为中心的本体工程:该本体围绕一组已定义化学问题所需的概念、属性、关系及可执行程序构建,而非作为化学的通用表示。所实现的产物结合了以JSON和RDF/Turtle序列化的轻量级本体,以及用于电子结构、周期趋势、氧化态、氧化物与氢化物行为及相关中学推理模式的确定性Python规则。一个独立的专家编码 fallback(回退机制)处理尚未被通用规则覆盖的问题类别。该系统在300个人类编写并经人工验证的化学问题上进行了测试:完整系统匹配了300个参考答案中的296个(准确率98.67%);本体驱动的规则子集覆盖269个问题,匹配266个参考答案(准确率98.88%);31个问题由特定任务的专家编码回退机制处理,其中30个匹配。由于同一组问题既用于本体构建又用于评估,这些结果衡量的是已实现的覆盖范围与内部一致性,而非独立泛化能力。我们分析了4个不匹配案例,区分了结构验证与化学正确性,并定义了未来架构:其中语言模型主要充当从用户语言到标准化本体任务框架的转换器。令牌效率被作为未来对照研究的可检验假设提出,而非当前工作的结果。

英文摘要

Large language models can interpret natural-language chemistry questions, but their internal reasoning is difficult to inspect, constrain, and validate. This paper presents ChemOntoRule, a proof-of-concept symbolic core for AI-assisted school-level chemistry problem solving. The central design choice is task-centric ontology engineering: the ontology is constructed around the concepts, properties, relations, and executable procedures required by a defined collection of chemistry problems, rather than as a universal representation of chemistry. The implemented artifact combines a lightweight ontology serialized in JSON and RDF/Turtle with deterministic Python rules for electronic structure, periodic trends, oxidation states, oxide and hydride behavior, and related school-level reasoning patterns. A separate expert-coded fallback handles problem families not yet represented by general rules. The system was examined on 300 human-authored and manually validated chemistry problems. The complete system matched 296 of 300 reference answers (98.67%). The ontology-driven rule subset covered 269 problems and matched 266 references (98.88%); 31 problems were handled by task-specific expert-coded fallbacks, with 30 matches. Because the same collection informed ontology construction and evaluation, these results measure implemented coverage and internal consistency, not independent generalization. We analyze the four mismatches, distinguish structural validation from chemical correctness, and define a future architecture in which a language model acts primarily as a translator from user language into a normalized ontological task frame. Token efficiency is presented as a testable hypothesis for future controlled studies, not as a result of the current work.

Comments15 pages, 4 figures, 5 tables. Proof-of-concept evaluation of a task-centric chemistry ontology and deterministic rule engine on 300 human-authored and manually validated school-level problems. The proposed LLM translation layer and token-efficiency hypothesis are not evaluated in the current study

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑