arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00354cs.CRcs.AIcs.CE

证明门控签名:在状态漂移下保持有效的求解器检查交易防护,用于链上AI智能体

Proof-Gated Signing: Solver-Checked Transaction Guards that Hold Under State Drift for Onchain AI Agents

Bravish Ghosh

AI总结:

针对链上AI智能体被诱导发起有害交易的问题,提出证明门控签名(PGS),通过SMT求解器验证交易效果并强制链上后置条件,在260个场景测试中阻止93.6%有害场景且零漂移收益。

AI中文摘要:

控制钱包的AI智能体可读取攻击者可触及的内容,因此可能被引导提出有害交易。通常的最后一道防线是预签名检查:静态白名单、LLM审查器或交易模拟。三者都存在一个共同缺陷:检查描述的是检查时的链状态,但交易在稍后的状态中执行,而该状态可能被对手通过前置运行、合约升级或代币参数变更来塑造。我们将此称为状态漂移。我们提出证明门控签名(PGS),它模拟拟议交易,提取其影响,并使用SMT求解器针对预言机不确定性区间内的每个价格检查声明式价值与权限策略。随后,它编译链上后置条件(钱包余额界限、收款人收据、授权上限和所有权),并证明满足这些条件的每次执行也满足策略。智能体的智能合约钱包以原子方式强制执行这些条件,因此该保证适用于任意漂移下已执行的交易。在一个包含260个场景的开放测试平台上(14个攻击家族,包括五个漂移和两个自适应家族,以及12个良性家族),危害从攻击者余额而非任何策略来衡量,PGS阻止了140个有害场景中的93.6%,并通过了97.5%的良性场景。仅模拟检查阻止了57.9%,静态白名单阻止了71.4%。在PGS下,50个漂移场景中没有一个产生攻击者收益。唯一未阻止的家族,即策略内抽取,受限于每会话预算。我们还发现,给LLM审查器一个干净的漂移前模拟使其更可能批准漂移攻击。开销约为每次检查41k gas和0.1-0.2秒。

英文摘要:

AI agents that control wallets read attacker-reachable content, so they can be steered into proposing harmful transactions. The usual last line of defense is a pre-signing check: a static allowlist, an LLM reviewer, or a transaction simulation. All three share a gap: the check describes the chain state at check time, but the transaction executes in a later state that an adversary can shape through front-running, contract upgrades or token-parameter changes. We call this state drift. We present Proof-Gated Signing (PGS), which simulates a proposed transaction, extracts its effects, and uses an SMT solver to check a declarative value-and-permission policy for every price in an oracle-uncertainty band. It then compiles on-chain post-conditions (wallet balance bounds, payee receipts, allowance caps and ownership) and proves that every execution satisfying them also satisfies the policy. The agent's smart-contract wallet enforces them atomically, so the guarantee applies to the executed transaction under arbitrary drift. On an open testbed of 260 scenarios (14 attack families including five drift and two adaptive families, and 12 benign families), with harm measured from attacker balances rather than from any policy, PGS prevented 93.6% of the 140 harmful scenarios and passed 97.5% of the benign ones. Simulation-only checking prevented 57.9% and a static allowlist 71.4%. None of the 50 drift scenarios produced attacker gain under PGS. The only unprevented family, an in-policy drain, was bounded by the per-session budget. We also find that giving an LLM reviewer a clean pre-drift simulation made it more likely to approve a drift attack. Overhead is about 41k gas and 0.1-0.2 s per check.

补充信息

↑