LLM数据智能体的追踪完整性:现实世界系统中可审计结构化推理的愿景
Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems
查看机构详情
- WAI USA Research Labs(WAI美国研究实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文针对LLM数据智能体提出追踪完整性标准,结合执行契约实现,通过CAIT率实验表明需同时评估答案准确率与可审计计算,为现实系统的LLM数据智能体提供评估方案。
中文摘要 AI 辅助
答案准确率并非LLM数据智能体的充分可靠性信号。在结构化数据任务中,基准正确的答案可能由无效追踪产生。本文提出追踪完整性(Trace Integrity),这是一种部署可靠性标准,用于评估答案背后记录的计算是否具备显式、可执行、符合模式、算子保真、可重放、答案一致及可审计的特性。我们将结构缺口(Structure Gap)确定为使追踪完整性成为必要的部署失效模式:自然语言推理与自由形式的理由无法可靠指定现实世界系统所需的算子级程序。我们通过执行契约(execution contracts)实现追踪完整性,执行契约是将用户意图与模式元素、算子计划、假设、可执行查询、验证状态及最终答案关联起来的结构化制品。我们还提出CAIT(正确答案/无效追踪)率,用于衡量仅答案评估将计算上无支撑的输出计为成功的频率。在BIRD Mini-Dev上的实证演示中,直接SQL、操作摘要+SQL、契约优先SQL的答案准确率分别为20%、22%、24%,而它们的追踪完整性通过率分别为39%、43%、40%,CAIT率仍高达55%、59.1%、45.8%,表明答案准确率、追踪有效性及静默失效风险是不同的评估信号。因此,现实世界的LLM数据智能体不仅应通过输出是否匹配参考答案来评估,还应通过输出是否有可审计的计算支撑来评估。
英文摘要
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. We identify the Structure Gap as the deployment failure mode that makes Trace Integrity necessary: natural-language reasoning and free-form rationales do not reliably specify the operator-level programs required by real-world systems. We operationalize Trace Integrity with execution contracts, structured artifacts that bind user intent to schema elements, operator plans, assumptions, executable queries, verification status, and final-answer linkage. We also introduce CAIT (Correct Answer / Invalid Trace) Rate, which measures how often answer-only evaluation counts computationally unsupported outputs as successes. In an empirical demonstration on BIRD Mini-Dev, Direct SQL, Operation Summary + SQL, and Contract-First SQL achieve answer accuracies of 20%, 22%, and 24%, while their Trace Integrity Pass Rates are 39%, 43%, and 40% and their CAIT Rates remain high at 55%, 59.1%, and 45.8%, showing that answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals. Real-world LLM data agents should, therefore, be evaluated not only by whether their outputs match a reference answer, but by whether those outputs are backed by auditable computation.