arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14121cs.AIcs.DBcs.IR

语义知识技术:语义网所忽视的,以及它从未拥有的

Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had

  • Glycan and Life Systems Integration Center (GaLSIC)(聚糖与生命系统整合中心(GaLSIC))
  • Soka University(创价大学)

机构由 AI 辅助整理,请以论文原文为准。

Achille Zappa

AI总结:

本文指出语义网标准遗漏了主张条件、操作基础和覆盖范围,提出语义知识技术方案,以五测试和七层架构定义理解,并引入大型知识模型等概念,作为可证伪的研究议程。

AI中文摘要:

语义网旨在为信息提供机器可解释的形式,以便软件能够对其进行整合和推理。其标准已成为科学知识基础设施,但它所承诺的机器能力并未实现,如今回答科学知识问题的系统是语言模型,它们对其所知内容没有可检查的说明。本文认为最初的目标是正确的,而技术方案不完整,指出了缺失之处,并将扩展方案命名为语义知识技术:将相同的技术核心从其网络发布起源中脱离出来,应用于任何地方所持有的知识。诊断结果是,这些标准将真值形式化,却遗漏了三件事:主张成立的条件、其术语允许的操作,以及基础覆盖范围的任何说明。没有条件,就无法判断矛盾和适用性;没有操作基础,持有陈述不赋予任何能力;没有声明的覆盖范围,系统无法识别自身内容的边界,而在开放世界假设下这是无法推断的。本文将“理解”一词固定为五个可测试的测试(检查、连接、推导、行动、界定),并提出了一个七层架构,其中前三层是使能层,其余层是它们所使能的认知能力。然后定义了该方案隐含的三个术语:大型知识模型,一种输出单元是对可寻址主张的引用而非标记的模型;SLKM,智能体从声明的来源为自己构建的知识库;以及语义通用人工智能,被表述为关于必要条件的可证伪立场,而非系统。一个分级阶梯取代了不可测试的“通用”一词。本文将其作为研究议程提出,并指出了其最薄弱点和反驳条件。

英文摘要:

The Semantic Web set out to give information a machine-interpretable form so that software could integrate and reason over it. Its standards became scientific knowledge infrastructure, but the machine competence it promised did not follow, and the systems now answering questions over scientific knowledge are language models holding no inspectable account of what they know. This paper argues the original goal was right and the technical programme incomplete, states what is missing, and names the extended programme Semantic Knowledge Technologies: the same technical core carried out of its web-publishing origin and applied to knowledge wherever held. The diagnosis is that the standards formalised truth while omitting three things: the conditions under which a claim holds, the operations its terms permit, and any account of what a base covers. Without conditions, contradiction and applicability cannot be judged; without operational grounding, holding a statement confers no ability; without declared coverage, a system cannot recognise the boundary of its own content, which under the open-world assumption cannot be inferred. The paper fixes the word understanding to five measurable tests (check, connect, derive, act, delimit) and sets out a seven-layer architecture in which the first three layers are enabling and the rest the cognitive capabilities they make possible. It then defines three terms the programme implies: Large Knowledge Model, a model whose unit of output is a reference to an addressable claim, not a token; SLKM, the knowledge base an agent builds for itself from declared sources; and Semantic Artificial General Intelligence, stated as a falsifiable position about necessary conditions, not a system. A graded ladder replaces the untestable word general. It is offered as a research agenda, with its weakest points and refutation condition named.

↑