arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可信的领域特定人工智能:面向结构化知识检索与推理

Trustworthy Domain-Specific AI for Structured Knowledge Retrieval and Reasoning

Ryan C. Barron

arXiv 2610.08894首次发表:更新:

AI 中文总结

本论文提出一种可扩展架构,将非结构化领域文本转化为结构化知识,通过Binary Bleed和HNMFk等方法及T-SRAG检索,提升检索精度并减少幻觉,应用于多个领域。

AI 中文摘要

本论文提出了一种可扩展的架构,用于将非结构化的领域特定文本转化为结构化知识,以支持检索与推理。该架构将半自动语料库构建、语义结构化、检索与推理整合为一个可解释的流水线。研究引入了Binary Bleed,一种改进的二分搜索方法,用于降低非负矩阵分解(NMF)中低秩搜索的复杂度;以及具有自动潜在特征选择的层次化NMF(HNMFk),一种深度自适应的主题建模方法,能够在领域专家的指导下生成可解释的分类体系。这些表示被填充到类型化知识图谱和语义对齐的向量存储中,其中包含提取的潜在特征,并通过事件驱动的底层机制进行同步。张量结构化检索增强生成(T-SRAG)动态地将查询路由到不同的检索路径。对比对齐将文档和查询的嵌入映射到层次化主题结构,以提高语义保真度并减少幻觉。在检索之外,基于张量的链接预测能够识别并补全知识图谱中缺失的链接,支持基于引文结构的推理。在网络安全、法律、材料科学和医疗保健领域的应用展示了在检索精度、早期趋势检测、假设生成和幻觉缓解方面的改进。本论文为可信的、领域特定的人工智能系统提供了一个可部署的、模块化的基础,这些系统能够对结构化知识进行检索和推理。

英文摘要

This dissertation presents a scalable architecture for transforming unstructured, domain-specific text into structured knowledge for retrieval and reasoning. It integrates semi-automatic corpus curation, semantic structuring, retrieval, and inference into an interpretable pipeline. The research introduces Binary Bleed, an adapted binary search method that reduces low-rank search complexity for Non-negative Matrix Factorization (NMF), and Hierarchical NMF with automatic latent feature selection (HNMFk), a depth-adaptive topic modeling method that produces interpretable taxonomies guided by subject matter experts. These representations populate a typed Knowledge Graph and a semantically aligned Vector Store containing extracted latent features, synchronized through an event-driven substrate. Tensor-Structured Retrieval-Augmented Generation (T-SRAG) dynamically routes queries across retrieval paths. Contrastive alignment maps document and query embeddings to hierarchical topic structures to improve semantic fidelity and reduce hallucinations. Beyond retrieval, tensor-based link prediction identifies and completes missing links in the Knowledge Graph, supporting inference grounded in citation structure. Applications across cybersecurity, law, materials science, and healthcare demonstrate improvements in retrieval precision, early trend detection, hypothesis generation, and hallucination mitigation. The dissertation provides a deployable, modular foundation for trustworthy, domain-specific AI systems that retrieve and reason over structured knowledge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑