arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Quipu:一种受管控的双时态知识图存储系统

Quipu: A Governed Bitemporal Knowledge Graph Store

Steve Brown

arXiv 2608.16813首次发表:更新:

AI 中文总结

研究针对智能体知识图存储的四项默认规则缺陷,提出受管控双时态知识图存储系统Quipu,经Census、DEMM-Bench等基准测试,其在缺陷检测、裁决准确性、管控问题回答等方面表现优于基线。

AI 中文摘要

智能体当前正在生成知识图,但知识图存储系统仍沿用人类策划时设定的默认规则:接受当前写入的数据、后续再清理、仅保留一个时间轴或不保留时间轴、将所有写入者的事实视为同等可信,且将管控工作交由仪表板和中间件处理。这四项默认规则各自便捷,但在智能体工作负载下共同运行时无法维持。我们提出Quipu,这是一种可嵌入的存储系统,它反转了上述四项默认规则:任何事实仅能通过一个门控进入,该门控的谓词会评估待处理的后状态;数据、信任标签、裁决以及规则本身均为双时态;命名图是权威与信任的单元,在格结构下构成,该格的一个不变量是构成永远不会扩大;管控规范Σ、追踪记录以及已签名的裁决均为其所管控存储系统中的事实,使得审计T ⊨ Σ成为一项查询。我们通过Census进行评估,Census是一个确定性多写入者生命周期,其单次种子运行会针对植入的真实值对每个研究问题进行评分:门控存储系统最终在6个植入缺陷中检测到0个,而非门控存储系统则检测到6个;所有7次构成探测均符合格契约;50个已满足的裁决在其时间点均能准确重新推导,而若采用仅最新规则集,所有50个裁决都会被错误报告;SARC参考检查器与存储内审计裁决完全一致,仅在覆盖语义上存在差异。来自受管控写入的记录追踪揭示了一个实时执行缺口,审计会指出该缺口及对应的补救措施。在DEMM-Bench这一外部决策证据充足性基准上,对导出记录的仅内容读取在全部8种退化条件下均正确回答了全部512个属性级管控问题,且无过度声明,而容器存在基线则在多达87.5%的问题上过度声明——且该运行还揭示了一个缺口,即拒绝裁决所证明的内容存在问题,我们随后修复了该缺口。

英文摘要

Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware. These four defaults are individually convenient and jointly untenable under agent workloads. We present Quipu, an embeddable store that inverts all four: no fact enters except through a gate whose predicates evaluate the pending post-state; data, trust labels, verdicts, and the rules themselves are bitemporal; named graphs are the unit of authority and trust, composed under a lattice whose one invariant is that composition never widens; and the governance specification $Σ$, the trace, and signed verdicts are facts in the store they govern, making the audit $T \models Σ$ a query. We evaluate with Census, a deterministic multi-writer lifecycle whose single seeded run scores every research question against planted ground truth: the gated store ends with 0 of 6 planted defects versus 6 of 6 ungated; all 7 composition probes uphold the lattice contract; 50 of 50 satisfied verdicts re-derive faithfully as of their instant while all 50 would be misreported under a latest-only rule set; and the SARC reference checker agrees with the in-store audit verdict-for-verdict, differing only on coverage semantics. A recorded trace from a governed writer surfaces a live enforcement gap the audit names with its remediation. On DEMM-Bench, an external decision-evidence sufficiency benchmark, a content-only reading of the exported records answers all 512 property-level governance questions correctly with zero overclaim under all eight degradation conditions, while container-presence baselines overclaim on up to 87.5% of them -- and the run surfaced, and led us to close, a gap in what a denial's verdict attests.

Comments15 pages, 2 figures, 4 tables. Subtitle: "Start strict: rethinking knowledge-graph defaults for agent-written knowledge". Source and the benchmark/census/ artifacts behind every reported number are archived at doi:10.5281/zenodo.21878428 (concept DOI, resolves to newest release). Development repository: github.com/scbrown/quipu

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑