基于OWL 2 DL的神经符号学习:通过基于后承的编译到可微电路
Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
浏览论文内容
中文总结 AI 辅助
该研究提出Baobab,将$\text{SROIQ}$本体编译为可微句子决策图,结合神经感知与本体推理,在MNIST任务上缓解非Horn描述逻辑中的推理捷径,实现贝叶斯最优后验。
中文摘要 AI 辅助
OWL 2 DL本体以描述逻辑$\boldsymbol{\text{SROIQ}}$为基础,在生物医学和语义网中表达大型知识库。针对描述逻辑的神经符号(NeSy)学习器要么将本体嵌入连续空间,放弃经典后承关系,要么局限于具有单一规范模型的Horn片段$\boldsymbol{\text{EL}^{++}}$。我们提出Baobab,它将具有有限ABox的$\text{SROIQ}$本体编译为句子决策图(SDD):在基于后承的演算下饱和命题核心,并在活跃域上实例化剩余的$\text{SROIQ}$特征(名目、数量限制和角色公理)。SDD的证据条件加权模型计数随后训练感知网络,在部分ABox监督下识别真实图像:在使用所有独特$\text{SROIQ}$特征的本体上,卷积神经网络(CNN)学习读取MNIST数字并结合后继关系,恢复出独立感知仅能达到随机水平的潜在本体概念。当监督允许多个与本体一致的补全时,独立感知会陷入某一个补全,这是一种推理捷径;我们证明,由查询的正当性索引的混合模型可以表示独立感知无法实现的校准后验,且从电路枚举的补全中初始化该混合模型,在真实图像MNIST任务上达到贝叶斯最优后验,而单加权模型计数(single-WMC)和学习到的混合模型(BEARS集成假设类)则无法做到:据我们所知,这是首次在非Horn描述逻辑中表征并缓解推理捷径。编译器的可靠性和表示结果在Lean 4中经过机器验证。代码可在this https URL获取。
英文摘要
OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment $\mathcal{EL}^{++}$, which has a single canonical model. We present Baobab, which compiles a $\mathcal{SROIQ}$ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence-based calculus and instantiates the remaining $\mathcal{SROIQ}$ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence-conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive $\mathcal{SROIQ}$ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology-consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes-optimal posterior on a real-image MNIST task where single-WMC and learned mixtures (the BEARS-ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non-Horn description logic. Soundness of the compiler and the representation result are machine-checked in Lean 4. Code is available at https://github.com/bio-ontology-research-group/baobab.
发表机构
- King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。