arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17107cs.AI

符号分离:将深度智能体锚定于知识图谱以实现可信运维数据分析

Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics

  • University of Bologna(博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Baibek Davletiyarov, Junaid Ahmed Khan, Andrea Bartolini

AI总结:

提出符号分离方法,通过本体约束的虚拟知识图谱与确定性验证,将深度智能体锚定于知识图谱,使神经符号深度分析师在49.9TB遥测数据上将任务成功率从43%提升至86%,并降低2.4倍令牌成本。

AI中文摘要:

生成式人工智能承诺为数据中心和工业4.0设施的海量数值遥测数据提供自然语言访问,然而文本到查询和工具使用智能体仍然不可靠:即便是前沿模型也只能正确回答一半以上的真实世界数据库问题,而涉及多步骤的运维问题则更少,因为大语言模型必须自行推断异构数据源之间的关系,并且会幻觉出这些关系,而不仅仅是字段。我们提出符号分离:一个深度智能体可以自由推理,但只能通过受本体约束的虚拟知识图谱并经过确定性的执行前验证来对数据采取行动。与工具API的接口契约不同,这种领域语义契约将一个复杂问题转化为一次经过验证的图遍历,而非由大语言模型推断的连接。该方法实例化为神经符号深度分析师,并在49.9 TB的超级计算机遥测数据上,与刚性工作流和非符号消融进行对比评估,将端到端任务成功率从43%提升至86%,防止了任何语法检查都无法捕获的静默数据完整性错误,并将令牌成本降低了2.4倍,使得较小的本地模型能够胜过较大的模型。

英文摘要:

Generative AI promises natural language access to the massive numerical telemetry of data centers and Industry 4.0 installations, yet text-to-query and tool-using agents stay unreliable: even frontier models answer little more than half of real-world database questions, and far fewer of the multi-step, operational ones, because the LLM must compose how heterogeneous sources relate and hallucinates the relations, not just the fields. We propose symbolic separation: a deep agent reasons freely but may act on data only through an ontology-constrained Virtual Knowledge Graph with deterministic pre-execution validation. Unlike a tool API's interface contract, this domain-semantic contract turns a complex question into one validated graph traversal instead of LLM-inferred joins. Instantiated as the Neurosymbolic Deep Analyst and evaluated on 49.9 TB of superconputer telemetry against a rigid workflow and a non-symbolic ablation, it raises end-to-end task success from 43% to 86%, prevents silent data-integrity errors that no syntactic check catches, and cuts token cost by 2.4x, letting a smaller on-premise model outperform a larger one.

↑