arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09286cs.DB

无痕迹,无主张:数据库智能体的两种契约

No Trace, No Claim: Two Contracts for Database Agents

Xiaofei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出数据库智能体所需的计划契约与证据契约,通过双时态图系统TGMS实现可验证执行与忠实主张,提升结果可靠性。

中文摘要 AI 辅助

大语言模型智能体能够生成数据库操作并解释其结果,但当前的接口通常在生成的计划、执行条件和向用户呈现的主张之间存在差距。我们认为,面向智能体的数据系统需要两种可执行的契约。计划契约定义了智能体可以执行和引用的内容;证据契约记录了结果支持主张所依据的信念状态、完整性和来源。我们在TGMS中实例化了这些契约,这是一个双时态图系统,其中大语言模型在固定的时间算子接口上进行规划。静态验证器在执行前检查计划,主张验证器对照内容寻址的痕迹检查类型化主张。实时模型运行暴露了仅输入模式和价值接地所遗漏的两种失败:不存在的结果字段和将页面局部计数报告为完整结果计数。结果字段检查使第一种失败成为可修复的拒绝。完整性传播在所有15个受控案例中检测到第二种失败,而在禁用时则全部遗漏。在冻结的CollegeMsg工作负载上,TGMS达到0.408的类型化答案准确率,而评估的基线为0.064至0.284。在Bitcoin-OTC上,对同一双时态存储的直接SQL与TGMS匹配,显示固定算子接口没有普遍准确性优势。在修正探针上,TGMS和双时态SQL回答历史信念问题,而最新状态基线无法回答。在主张门控之前,220个答案中有21个包含不支持的门控主张;门控后,199个发出的答案中没有一个包含,代价是少21个答案。计划契约使无效计划可拒绝,并在记录条件下使执行可重现,而证据契约使门控主张忠实于引用的证据。两者都不保证对用户意图的正确解释。

英文摘要

LLM agents can generate database operations and explain their results, but current interfaces often leave a gap between generated plans, execution conditions, and claims presented to users. We argue that agent-facing data systems need two enforceable contracts. A plan contract defines what an agent may execute and reference; an evidence contract records the belief state, completeness, and provenance under which a result supports a claim. We instantiate these contracts in TGMS, a bi-temporal graph system in which an LLM plans over a fixed temporal operator interface. A static verifier checks plans before execution, and a claim verifier checks typed claims against content-addressed traces. Live model runs exposed two failures missed by input-only schemas and value-only grounding: nonexistent result fields and page-local counts reported as complete-result counts. Result-field checking makes the first a repairable rejection. Completeness propagation detects the second in all 15 controlled cases and misses all 15 when disabled. On the frozen CollegeMsg workload, TGMS reaches 0.408 typed-answer accuracy versus 0.064--0.284 for the evaluated baselines. On Bitcoin-OTC, direct SQL over the same bi-temporal store matches TGMS, showing no universal accuracy advantage for the fixed operator interface. On correction probes, TGMS and bi-temporal SQL answer historical-belief questions, while latest-state baselines cannot. Before claim gating, 21 of 220 answers contain an unsupported gated claim; after gating, none of 199 emitted answers does, at the cost of 21 fewer answers. The plan contract makes invalid plans rejectable and execution reproducible under recorded conditions, while the evidence contract makes gated claims faithful to cited evidence. Neither guarantees correct interpretation of user intent.

发表机构

  • University of Memphis(孟菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑