LEDGER:用于审计大语言模型智能体的从主张到证据的追踪图
LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents
- Carnegie Mellon University(卡内基梅隆大学)
- Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM智能体输出审计瓶颈,提出LEDGER系统构建分层追踪图,通过分组节点、添加语义边等实现以证据为中心的审计,揭示工作流决策等关键信息。
AI中文摘要:
大语言模型(LLM)智能体如今可执行涉及复杂工具使用、代码执行、文件编辑及生成工件的长周期技术工作流。随着智能体工作效率提升,生产力瓶颈从生成输出转向审计这些输出是否正确且可信。智能体可观测性系统能呈现细粒度执行事件,但仅靠可观测性仍需审阅者重构哪些动作、工件及验证步骤对特定结论重要。我们提出LEDGER——用于执行审阅的分层证据与决策图(Layered Evidence and Decision Graphs for Execution Review),这是一种追踪与审阅系统,可在观测到的智能体会话上构建分层追踪图。LEDGER保留追踪记录,同时将其分组为证据节点与工作流节点,将工件表示为证据锚点,并添加将主张与支撑动作、工件及检查连接的类型化语义边。通过数据分析与编码示例,我们展示所得追踪如何为以证据为中心的审计揭示工作流决策、工件谱系、修复步骤、验证覆盖范围及主张支撑路径。
英文摘要:
Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits, and generated artifacts. As agents do more work faster, the productivity bottleneck shifts from producing outputs to auditing whether those outputs are correct and trustworthy. Agent observability systems make fine-grained execution events visible, but visibility alone still leaves reviewers to reconstruct which actions, artifacts, and validation steps matter for a particular conclusion. We introduce LEDGER - Layered Evidence and Decision Graphs for Execution Review, a tracing and review system that builds layered trace graphs over observed agent sessions. LEDGER preserves Trace Records while grouping them into Evidence Nodes and Workflow Nodes, representing artifacts as evidence anchors, and adding typed semantic edges that connect claims to supporting actions, artifacts, and checks. Through data-analysis and coding examples, we show how the resulting traces expose workflow decisions, artifact lineage, repair steps, validation coverage, and claim-support paths for evidence-centered audit.