arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Janus:面向智能体LLM的“证据先于效果”Saga与离线可验证溯源

Janus: Evidence-Before-Effect Sagas and Offline-Verifiable Provenance for Agentic LLMs

Mustafa Arslan

arXiv 2609.38266首次发表:更新:

发表机构

Independent Researcher Istanbul, T\"urkiye

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Janus通过签名哈希链日志在效果释放前记录智能体LLM的提议与裁决,实现离线可验证溯源,并在借贷实验中阻止了超出授权的支付,同时发现并报告了批准后替换漏洞。

AI 中文摘要

智能体大语言模型(LLM)如今能够通过工具转移资金,然而对其行为的记录通常只是模型自身过程在效果产生之余生成的痕迹。Janus将记录置于效果路径之上。一个步骤的提议、对其的裁决以及来自验证器或人工的任何回答,在该步骤得以运行或其效果得以释放之前,均以签名、哈希链日志的形式持久化;在声明密钥后,每个回答均由给出或转发它的人签名。门控是日志的纯函数,审计员可离线地从日志和单个公钥重新推导出每一项裁决。在MCP边缘,效果被持有直至上述过程完成;而通过SDK(我们的模型实验所用),协作客户端仅在之后才运行该效果。我们通过崩溃注入(进程内144次杀死,通过守护进程81次)、离线验证一个包含1亿事件的日志(254.5秒),以及在借贷工作流后端运行真实模型(对相同的记录模型输出分别进行受治理运行和普通运行)来评估Janus。当借贷授权包含在模型提示中时,对比结果为0对0。当授权仅存在于策略中且金额单位已声明时,模型批准了六笔超出授权的贷款,其中三笔无注入(一次运行还丢弃了单位,批准了三笔);普通智能体支付了全部六笔,而Janus未支付任何一笔,每笔均被确定性验证器拒绝且可离线重新推导。对记录输入使用始终批准的预言机,得到20笔和21笔超出授权的声明(对比0),尽管Janus支付了四笔和三笔,其声明金额低于请求金额。设计实验暴露了一个通过自身审计的系统中的批准后替换实例:一个针对某次尝试的批准被计入另一个提议,将人工批准从100移至1,000,000。我们报告了该问题、首个修复方案及绕过它的五种途径,以及Janus不保证的内容。

英文摘要

Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. Janus puts the record on the effect path. A step's proposal, the verdict on it and any answer from a validator or a person are durable in a signed, hash-chained log before the step may run or its effect be released; with keys declared, each answer is signed by whoever gave or relayed it. Gates are pure functions of that log, and an auditor re-derives every verdict offline from the log and one public key. At the MCP edge the effect is held until then; through the SDK, which our model experiment uses, a cooperating client runs it only afterwards. We evaluate Janus under crash injection (144 kills in-process, 81 through the daemon), by verifying a 100-million-event log offline (254.5 s), and with a real model behind a lending workflow, run governed and plain on the same recorded model outputs. With the lending mandate in the model's prompt, the comparison was 0 against 0. With it only in the policy and the amount's unit stated, the model approved six loans declared over the mandate, three with no injection (a run that also dropped the unit approved three); the plain agent paid all six and Janus none, each refused by a deterministic validator and re-derivable offline. An always-approve oracle over the recorded intakes gave 20 and 21 declared over the mandate against 0, though Janus paid four and three whose declared amount understated the request. Designing the experiment exposed, in a system that passed its own audit, an instance of post-approval substitution: an approval keyed to an attempt was counted for a different proposal, moving a person's approval from 100 to 1,000,000. We report it, a first fix and the five routes around it, and what Janus does not guarantee.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑