一种证据优先的多大语言模型框架,用于可审计的关键基础设施依赖建模
An Evidence-First Multi-LLM Framework for Auditable Critical-Infrastructure Dependency Modeling
浏览论文内容
中文总结 AI 辅助
提出一种证据优先的多LLM框架,通过分离验证与融合步骤,从异构文档构建可审计的基础设施依赖模型,实验表明实体恢复优于依赖恢复,端点解析是主要瓶颈。
中文摘要 AI 辅助
关键基础设施的知识分布在异构、不完整且弱结构化的证据中,这使得依赖模型的自动构建变得困难且难以信任。大型语言模型(LLMs)可以从此类证据中提取结构化知识,但直接从LLM到图的生成存在不支持的关系、不一致的术语、错误的实体身份以及错误的依赖端点的风险。我们提出了一种证据优先的多LLM框架,用于从异构基础设施文档中构建基础设施知识库(IKBs)和基础设施依赖图(IDGs)。多个开放权重LLM独立地从规范化证据中提取候选实体和依赖关系,之后框架将证据验证、本体论接地、实体解析、依赖对齐、验证、融合和人工审查分开。证据支持、本体协调、端点解析、模型一致性和人工验证保持为不同的状态,而来源和未解决案例在整个过程中得以保留。经过验证的IKB随后被确定性地投影到IDG中,而不引入新的LLM生成的知识。在九个基础设施项目中的评估显示,实体恢复的召回率显著高于完整的有向依赖恢复,且规范端点解析是依赖构建中的主要约束。依赖关系的跨模型重叠也远低于实体,这表明模型通常产生非重叠的候选断言,而非稳定的多数共识。这些发现支持一个可审计的证据到IKB再到IDG的过程,其中不确定性被保留并逐步解决,而不是被压缩为单一的置信度或投票决策。
英文摘要
Critical-infrastructure knowledge is distributed across heterogeneous, incomplete, and weakly structured evidence, making dependency models difficult to construct automatically and difficult to trust. Large language models (LLMs) can extract structured knowledge from such evidence, but direct LLM-to-graph generation risks unsupported relationships, inconsistent terminology, incorrect entity identities, and erroneous dependency endpoints. We present an evidence-first multi-LLM framework for constructing Infrastructure Knowledge Bases (IKBs) and Infrastructure Dependency Graphs (IDGs) from heterogeneous infrastructure documentation. Multiple open-weight LLMs independently extract candidate entities and dependencies from normalized evidence, after which the framework separates evidence verification, ontology grounding, entity resolution, dependency alignment, validation, fusion, and human review. Evidence support, ontology reconciliation, endpoint resolution, model agreement, and human validation remain distinct states, while provenance and unresolved cases are preserved throughout. The validated IKB is then projected deterministically into the IDG without introducing new LLM-generated knowledge. Evaluation in nine infrastructure projects shows that entity recovery achieves substantially higher recall than complete directed dependency recovery and that canonical endpoint resolution is a major constraint in dependency construction. Cross-model overlap is also much lower for dependencies than for entities, indicating that the models often produce non-overlapping candidate assertions rather than a stable majority consensus. These findings support an auditable evidence-to-IKB-to-IDG process in which uncertainty is preserved and resolved progressively rather than collapsed into a single confidence or voting decision.
发表机构
- Louisiana State University(路易斯安那州立大学)
机构由 AI 辅助整理,请以论文原文为准。