发表机构
University of Chinese Academy of Sciences; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院大学; 中国科学院信息工程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具智能体声明与执行证据脱节的问题,提出声明锚定的执行契约,联合绑定声明、来源、执行顺序及版本,实现可重放的审计,在跨对象攻击中检测率达0.9961。
AI 中文摘要
使用工具的人工智能体可以展示引用和执行日志,但留下了一个关键的关联未被审计:即展示给用户的声明是否确实是由已提交的执行所发出的,并得到所引用来源的支持。因此,一个有效的引用和一个有效的执行轨迹可以在各自保持格式良好的同时,被移植到不同的声明、操作、运行或来源版本中。我们提出了一种声明锚定的执行契约,它联合绑定了发出的声明、其精确的来源片段、产生该声明的有序执行前缀,以及该执行所观察到的来源版本和访问状态。每个收据包含一个发射锚点,它能确定性地在已提交的答案或携带声明的操作中定位该声明,同时包含来源标识符、偏移量、哈希值、引文,以及一个域分离的执行承诺。一个确定性的完整性验证器在语义或任务标签被加入之前重建这些绑定。我们将这一完整性平面与一个可插拔的支持平面分离,因此结构有效性不会被用作蕴含关系的代理。该契约暴露了七个可独立测试的属性:声明-发射绑定、来源绑定、有序执行绑定、预言机分离、持久化对象重放、执行重运行一致性,以及版本/访问绑定。在1,280次跨对象攻击中,联合契约检测到1,275次替换(0.9961)。移除一个目标属性会将其攻击检测率降至0.0156-0.0625。在一个独立裁决的384对分割中,冲突感知的支持防护达到F1分数0.8865,误接受率为0.0729;在未见过的失败族上,这些比率分别为0.8679和0.0938。
英文摘要
Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid citation and a valid trace can therefore remain individually well formed while being transplanted across claims, actions, runs, or source versions. We introduce a claim-anchored execution contract that jointly binds the emitted claim, its exact source span, the ordered execution prefix that produced it, and the source version and access state observed by that execution. Each receipt contains an emission anchor that deterministically locates the claim inside a committed answer or claim-bearing action, together with source identifiers, offsets, hashes, quotes, and a domain-separated execution commitment. A deterministic integrity verifier reconstructs these bindings before semantic or task labels are joined. We separate this integrity plane from a pluggable support plane, so structural validity is not used as a proxy for entailment. The contract exposes seven independently testable properties: claim-emission binding, source binding, ordered-execution binding, oracle separation, persisted-object replay, execution-rerun consistency, and version/access binding. Across 1,280 cross-object attacks, the joint contract detects 1,275 substitutions (0.9961). Removing a targeted property reduces its attack-detection rate to 0.0156-0.0625. On an independently adjudicated 384-pair split, the conflict-aware support guard reaches F1 0.8865 and false acceptance 0.0729; on unseen failure families, these rates are 0.8679 and 0.0938.
Comments35 pages, 8 figures