发表机构
Syracuse University; HRL Laboratories(雪城大学; HRL实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EntailLLM通过逻辑编程结合领域知识验证LLM生成的漏洞发现路径,在三类CWE、四个LLM等多组实验中,将合并蕴含率从78%提升至98%,可端到端部署且无需逐设备调优。
AI 中文摘要
大语言模型(LLM)越来越多地被用于对软件漏洞进行推理,但其输出可能会悄然违反领域知识,这限制了它们在医疗设备等安全关键场景中的可靠性。现有研究要么将其输出视为需要评分的预测,要么将其限制在单个知识图谱内的路径中,两者均未检查对二进制文件的推理是否与独立的领域知识体系一致。我们提出了EntailLLM,它通过蕴含关系验证每个LLM提出的分析路径:该路径是二进制文件函数调用图的遍历,领域知识表示在单独的图谱中,验证在时间标注逻辑下对齐两者。在三类常见弱点枚举(CWE)类别、四个LLM、三种提示策略以及七个函数调用图节点规模从405到12696不等的二进制文件上,领域知识将合并蕴含率从78%提升至98%,仅在3%的实验中蕴含率出现下降。EntailLLM被端到端部署在真实医疗设备二进制文件上,无需针对每个设备进行调整即可达到98%的合并蕴含率。我们的系统继承了广义标注逻辑的形式保证,提供了既具有可解释性又基于明确定义语义的LLM输出逻辑验证。
英文摘要
Large language models are increasingly used to reason about software vulnerabilities, but their outputs can silently violate domain knowledge, limiting their reliability in safety-critical settings such as medical devices. Prior work either treats that output as a prediction to be scored or constrains it to walks within a single knowledge graph; neither checks whether reasoning over a binary is consistent with an independent body of domain knowledge. We present EntailLLM, which validates each LLM-proposed analyst path by entailment: the path is a traversal of the binary's function call graph, the domain knowledge is represented in a separate graph, and verification aligns the two under temporal annotated logic. Across three CWE classes, four LLMs, three prompting strategies, and seven binaries varying in size from 405 to 12,696 function call-graph nodes, domain knowledge raises pooled entailment from 78% to 98%, with entailment decreasing in only 3% of the experiments. EntailLLM is deployed end-to-end on real medical-device binaries, reaching 98% pooled entailment without per-device tuning. Our system inherits the formal guarantees of generalized annotated logic, providing logical verification of LLM output that is both explainable and grounded in well-defined semantics.