arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17220cs.CRcs.AI

PACE:用于去中心化金融中安全AI智能体的策略验证合约执行

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

Rabimba Karanjai, Yang Lu, Richard Williamson, Hemanth Hm, Prakhar Mehrotra, Lei Xu, Weidong, Shi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出PACE框架,通过引入类型化交易意图等机制,实现去中心化金融中AI智能体的安全交易授权,在确定性沙箱中实现0.00不安全执行率和0.00误报率,提升了DeFi场景下AI智能体的安全性。

中文摘要 AI 辅助

自主AI智能体正逐渐成为去中心化金融(DeFi)相关操作的接口,例如兑换、借贷操作及收益管理。由于这些智能体依赖大型语言模型(LLM)规划交易,因此继承了LLM易受提示注入攻击的特性,且缺乏将验证者的批准与最终提交至链上的具体交易绑定的机制。我们提出PACE(Policy-Attested Contract Execution,策略验证合约执行),这是一种交易级授权框架,介于基于LLM的智能体与链上执行之间。PACE引入了类型化交易意图、确定性策略验证器,以及已签名的策略决策记录(PDRs),这些记录通过密码学方式将已批准的意图、策略和模拟报告与具体的执行字节绑定,并具备重放保护和过期保护功能。一个Solidity智能合约账户在链上强制执行PDR签名,测量的 gas 开销为29,826-31,822。我们在涵盖四个攻击类别及良性实用场景的40项任务上,将PACE与六个基线进行评估(共2,800次试验,10个随机种子)。在我们的确定性沙箱中,PACE在良性任务上实现了0.00的不安全执行率和0.00的误报率,而无防护基线的对应值为0.80。消融研究表明,宽松的策略设置(提升57.5个百分点)和已接触合约白名单(提升12.5个百分点)是主要的安全组件。为测试相同的确定性下限是否适用于真实模型输出,该工件还提供了在完整任务套件上针对三个模型的实时LLM评估,且包含重复运行。主网分叉工具包用于存档RPC部署,但仅在生成对应工件时才报告分叉结果。这些辅助研究独立于确定性基准,且绝不替代确定性基准。我们将主张界定为可复现基准内的逻辑级安全,而非可部署的DeFi安全。

英文摘要

Autonomous AI agents are emerging as interfaces for decentralized finance (DeFi) actions such as swaps, lending operations, and yield management. Because these agents rely on large language models (LLMs) to plan transactions, they inherit the LLM's susceptibility to prompt injection and lack of mechanisms to bind a verifier's approval to the exact transaction ultimately submitted on-chain. We present PACE (Policy-Attested Contract Execution), a transaction-level authorization framework that interposes between an LLM-based agent and on-chain execution. PACE introduces typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind the approved intent, policy, and simulation report to the exact execution bytes, with replay and expiration protection. A Solidity smart account enforces PDR signatures on-chain with a measured overhead of 29,826-31,822 gas. We evaluate PACE against six baselines on 40 tasks spanning four attack categories plus benign utility (2,800 trials, 10 seeds). In our deterministic sandbox, PACE achieves a 0.00 unsafe execution rate and 0.00 false-positive rate on benign tasks, compared to 0.80 for the unguarded baseline. Ablation studies identify permissive policy settings (+57.5 pp) and the touched-contract allowlist (+12.5 pp) as the dominant safety components. To test whether the same deterministic floor holds for real model outputs, the artifact additionally provides a three-model live-LLM evaluation over the full task suite with repeated runs. A mainnet-fork harness is included for archive-RPC deployments, but fork results are reported only when the corresponding artifacts are generated. These auxiliary studies are separate from, and never substitute for, the deterministic benchmark. We frame our claims as logic-level safety within a reproducible benchmark rather than deployment-ready DeFi security.

发表机构

  • University of Houston(休斯顿大学)
  • PayPal(贝宝)

机构由 AI 辅助整理,请以论文原文为准。

↑