AI 中文总结
该研究提出基于POLANYI++的DARKSIDE连贯性审计方法,以OWL2格式的扩展知识图谱为基础,通过凭证轴分类指称对象并设置升级规则,在BSBench上验证其可缩小LLM的模式与路径结构差距,提升对无意义输入的检测能力。
AI 中文摘要
大型语言模型(LLMs)能够识别模式,但本身无法追踪连贯话语所需的排除路径。当输入基于伪造的权威、误用的机制或隐蔽的类比时,未受引导的LLM往往会将其视为有依据的内容,并将这种错误固化到生成的任何结构化输出中。带有POLANYI++的逻辑增强生成(LAG)是一种LLM引导方法,它使用启发式方法、本体论和问题解决方法进行隐性知识提取,以OWL2格式生成扩展知识图谱(XKG),但继承了同样的漏洞:复杂的无意义输入会与合法三元组一起被固化到图谱中,由于XKG是基于错误假设共同生成的,自动推理器几乎无法检测到。我们引入DARKSIDE,这是一种基于POLANYI++的连贯性审计方法。它将该路径形式化为随话语时间累积排除项的显式数据结构,并补充了一个凭证轴,将每个命名指称对象分类为有凭证(Warranted)、未证实(Unattested)、错误归属(Misattributed)或伪造(Fabricated),并设置了升级规则:当伪造率为正或未支持率超过阈值时,将委托风险评估(DelegationRiskAssessment)提升为不安全(UNSAFE)。我们在BSBench上评估了作为引导层的DARKSIDE,BSBench是一个包含软件工程、金融、医疗保健、物理和法律领域的100项复杂无意义内容的对抗性语料库,使用Claude Sonnet 4.6作为独立评判者。实证证据支持一个架构主张:当LLM的前向传递被本体介导的负路径装置包裹时,模式与路径之间的结构差距可以部分搭建。XKG充当缺失的记忆,而凭证轴则充当认知防火墙。
英文摘要
Large Language Models (LLMs) do not natively track the path of exclusions that a coherent discourse demands. When an input rests on a fabricated authority, a misapplied mechanism, or a surreptitious analogy, an unsteered LLM tends to engage with it as if it were well-posed, and this affects its generation. POLANYI++, an LLM-steering method that uses heuristics, ontologies and problem-solving methods for tacit-knowledge extraction, produces an Extended Knowledge Graph (XKG) in OWL2, but when a sophisticated nonsensical input is reified into the graph alongside the legitimate triples, it gets hardly detectable by automated reasoners, since the XKG is generated jointly with the wrong assumptions. We introduce DARKSIDE, a coherence-auditing method on top of POLANYI++. DARKSIDE formalises an explicit data structure of accumulated exclusions over discourse time, complemented by a warrant axis that classifies each named referent as Warranted, Unattested, Misattributed or Fabricated. The method is anchored in nine theoretical fragments unified under a shared deep frame of path integrity. The resulting DARKPOLANYI is evaluated as a steering layer over Gemini 3 on BSBench, a 100-item adversarial corpus of sophisticated-sounding nonsense across multiple domains, with Claude Sonnet 4.6 as an independent judge. DARKPOLANYI scores 1.89/2 mean versus 0.95/2 for the unsteered Gemini 3 Pro baseline; on the 97 cases with valid judgments in both arms, paired McNemar gives a paired bootstrap mean-diff = +0.92 (95% CI [+0.75, +1.08], p = 0.0001). The evidence supports an architectural claim: when an LLM forward pass is wrapped in an ontology-mediated auditing, structurally inevitable hallucination can be partially recovered. The XKG functions as the missing memory that LLMs lack, and the warrant axis as an epistemic firewall.
Comments23 pages, 4 figures, 5 tables