arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PolicyKG:一种用于将机构政策转换为SHACL知识图谱的智能体大语言模型流水线

PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs

Ponkrit Kaewsawee, Chaklam Silpasuwanchai, Chutiporn Anutariya

arXiv 2608.09028首次发表:更新:

发表机构

Asian Institute of Technology(亚洲理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出PolicyKG流水线,可自动将机构政策转换为SHACL知识图谱,在AIT语料库上取得高分类准确率,适配新领域仅需替换注册表,解决了人工转换的效率问题。

AI 中文摘要

机构政策以自然语言呈现,而合规性检查系统需要机器可读的约束条件,弥合这一差距目前仍依赖人工完成,PolicyKG则闭环解决该问题。它是一个大语言模型流水线,可读取政策PDF,将每个句子分类为义务、许可或禁止,将标签提升为一阶道义逻辑,并输出SHACL约束。四个阶段在LangGraph状态机上运行,每个阶段都有验证器。最关键的部分是语料库适配器:一个YAML词汇注册表,用于将大语言模型谓词锚定到目标本体,重定向到新领域只需替换注册表,无需重新训练模型。在亚洲理工学院(AIT)的政策与程序语料库(1663个句子,443条规则)上,PolicyKG的道义分类准确率达86.9%(Cohen's kappa=0.709)。三名标注员对50个样本独立重新标注,Fleiss' kappa=0.844。在69个形状子集上的SHACL形状正确性F1值为0.866。一阶逻辑(FOL)路径处理79.2%的规则,其余规则通过直接的自然语言到SHACL的回退路径处理。我们对全部443条规则进行了二阶及更高阶构造的审核,自动正则表达式检查表未标记出此类构造,第一作者对92个FOL回退案例的审核也确认了这一点,真实高阶逻辑(HOL)率的精确95% Clopper-Pearson上限为0.67%,这是针对一个语料库的审核发现,并非机构政策适用FOL的证明。将AIT注册表替换为通用数据保护条例(GDPR)注册表后,属性对齐的精确数量从1/15提升至11/15(Fisher精确检验p<0.001;Cohen's h=1.53)。在LexDeMod租赁合约基准(N=200)上,宏观F1降至0.370,因为租赁文本用“shall be entitled”表示许可——这正是注册表替换旨在解决的词汇不匹配问题。重复运行会产生哈希值相同的SHACL输出。

英文摘要

Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that gap is still done by hand. PolicyKG closes the loop. It is an LLM pipeline that reads a policy PDF, classifies each sentence as an obligation, permission, or prohibition, lifts the label into first-order deontic logic, and emits SHACL constraints. Four stages run on a LangGraph state machine with per-stage validators. The piece that matters most is the Corpus Adapter: a YAML vocabulary registry that grounds LLM predicates in a target ontology. Retargeting to a new domain means swapping the registry, not retraining a model. On the Asian Institute of Technology Policies and Procedures corpus (1,663 sentences, 443 rules), PolicyKG reaches 86.9% deontic classification accuracy (Cohen's kappa = .709). Three annotators independently re-label a 50-item sample and agree at Fleiss' kappa = .844. SHACL shape correctness on a 69-shape subset is F1 = .866. The FOL path handles 79.2% of rules; the rest go through a direct NL-to-SHACL fallback. We audited every one of the 443 rules for second- or higher-order constructs. An automated regex checklist flagged none, and a first-author pass on the 92 FOL-fallback cases confirmed the same. The exact upper 95% Clopper-Pearson bound on the true HOL rate is 0.67%. This is an audit finding for one corpus, not a proof of FOL sufficiency for institutional policy. Swapping the AIT registry for a GDPR registry raises exact property alignment from 1/15 to 11/15 (Fisher's exact p < .001; Cohen's h = 1.53). On the LexDeMod lease-contract benchmark (N = 200), Macro F1 drops to .370 because lease English uses "shall be entitled" for permission -- exactly the vocabulary mismatch registry swap is meant to fix. Repeated runs produce hash-identical SHACL outputs.

Comments23 pages, 3 figures, 4 tables. Under review at IJCKG 2026, Bangkok, Thailand

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑