Tracekit:面向自主编码智能体的防篡改意图-推理-行动审计系统
Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents
AI总结:
本研究提出开源无依赖的Tracekit系统,通过哈希链式账本捕获自主编码智能体的意图、推理与行动信息以实现防篡改审计,经多组实验验证了其篡改检测能力、性能表现与审计有效性。
AI中文摘要:
自主编码智能体会读取不可信文件、运行shell命令并在极少监管下生成子智能体,但其运行记录通常是可编辑的日志。我们提出Tracekit,这是一个无依赖的开源系统,可为每个智能体会话捕获三类信息通道:人类的指令内容(意图)、模型对自身推理的表述(自报告),以及它实际执行的操作(行动)。这些信息会被写入一条采用哈希链式结构、支持外部锚定的账本中,并进行交叉校验。Tracekit可接入Claude Code的生命周期事件,重构多智能体层级结构,通过预执行策略对工具调用进行管控,支持通过SDK或HTTP API接收来自其他智能体的事件,并提供一个可在浏览器中重新校验账本的实时观测界面。\n我们通过五项实验评估Tracekit的性能:(1)在1600次随机篡改测试中,该哈希链可检测出所有编辑、删除、重排序、伪造插入和断写操作;尾部截断和完整重链攻击仅能通过锚定机制检测,当锚定间隔为300条记录时,检测率降至0.47,与闭式模型的预测结果一致。(2)单个钩子的中位耗时为23.9毫秒,在账本记录数达到100000条时耗时仍保持平稳,且在16个并发写入者的场景下哈希链仍能保持正确性。(3)正则表达式管控门仅能拦截44次有害工具调用中的18次(占比41%),同时会错误拦截40次良性调用中的3次;简单的改写即可绕过该管控。(4)在14次真实的Claude Code运行中,智能体从未执行4处植入的间接提示注入攻击,且每次都对其进行了披露;提供商 withheld 了全部27个思考块的文本,因此自报告仅局限于可见的文本内容。(5)采用种子故障拼接(一种将隐蔽的错位步骤插入真实轨迹的新方法),在63条轨迹和126次审核调用中,规则标记检测出了49条故障轨迹中的40条(占比82%),独立大语言模型审核员检测出了全部98条故障轨迹(占比100%),针对未修改轨迹的假阳性率为1/14。我们已开源该系统、测试框架及所有轨迹数据。
英文摘要:
Autonomous coding agents read untrusted files, run shell commands and spawn sub-agents with little supervision, yet their record is usually an editable log. We present Tracekit, an open-source, dependency-free system that captures three channels for every agent session: what the human asked (intent), what the model said of its reasoning (self-report), and what it actually executed (actions). These are written to a hash-chained, externally anchorable ledger and cross-checked. Tracekit hooks into Claude Code's lifecycle events, reconstructs multi-agent hierarchies, gates tool calls with a pre-execution policy, accepts events from other agents via an SDK or HTTP API, and renders a live observer that re-verifies the ledger in the browser. We evaluate Tracekit in five experiments. (1) Across 1,600 random mutations, the chain detects every edit, deletion, reordering, forged insertion and torn write; tail truncation and full re-chaining are caught only by anchors, with detection falling to 0.47 at an anchoring interval of 300 records, matching a closed-form model. (2) A hook costs 23.9 ms median, flat up to 100,000 ledger records, and the chain stays correct under 16 concurrent writers. (3) A regular-expression gate blocks only 18 of 44 harmful tool calls (41%) while wrongly blocking 3 of 40 benign ones; trivial rewrites evade it. (4) In 14 real Claude Code runs, the agent never acted on four planted indirect prompt injections and disclosed each. The provider withheld the text of all 27 thinking blocks, so self-report was limited to visible prose. (5) With seeded-fault splicing, a new method that inserts concealed misaligned steps into real traces, over 63 traces and 126 reviewer calls, rule flags caught 40/49 (82%) of faulted traces and an independent LLM reviewer caught 98/98 (100%), with a false-positive rate of 1/14 on unmodified traces. We release the system, harness and all traces.