AI 中文总结
该研究针对编排不可逆状态转换的智能体,提出四维形式主义与七个安全不变量,在公共账本场景下验证其可提升智能体安全栈通过率,且可推广至其他不可逆状态交互场景。
AI 中文摘要
自主智能体越来越多地被要求对外部系统产生不可逆影响,例如转移资金、写入持久存储、驱动硬件。现有的智能体框架(ReAct、Reflexion、MCP)在基准测试中优化任务成功度,却很少关注不可逆副作用的安全性。我们将此类场景(即跨公共账本的价值转移)形式化为由(钱包、链、地址、协议)索引的四维空间中的状态转换,并利用该形式主义来陈述和证明我们称为执行保真度的保证:在存在规划器映射错误、结果模糊、重试、至少一次传递以及委托非人类调用者的故障模型下,会话在账本上实现的效果要么完全没有,要么恰好是向用户呈现的转换,且仅一次。该定理故意不声称呈现的转换与用户意图匹配——没有运行时层可以做出该判断——但它将这个无边界问题限制为对有限对象的单个谓词,这使得预览成为充分的控制而非形式主义。从保真度条件而非经验枚举中导出的七个安全不变量释放了该保证。实证上,在受控的N=60对抗性套件中,该栈在两个写激进的后备模型上比 naive-ReAct 基准提升了约74个百分点的通过率,但在写谨慎的模型上仅提升了约3个百分点——这表明对智能体安全栈的单模型评估几乎无法证伪。该系统已部署;跨8条链的108个生产写操作支持了故障分类。尽管评估场景是公共账本,但该形式主义和不变量适用于对不可逆外部状态采取行动的任何概率智能体。
英文摘要
Autonomous agents are increasingly asked to produce irreversible effects on external systems - transferring funds, writing to durable storage, actuating hardware. Existing agent frameworks (ReAct, Reflexion, MCP) optimize task success on benchmarks and give little attention to the safety of irreversible side-effects. We formalize one such setting, movement of value across public ledgers, as state transitions in a four-dimensional space indexed by (wallet, chain, address, protocol), and use that formalism to state and prove a guarantee we call execution fidelity: under a fault model admitting planner mis-mapping, ambiguous outcomes, retries, at-least-once delivery, and delegated non-human callers, a session's realized effect on the ledger is either nothing at all or exactly the transition that was rendered to the user, exactly once. The theorem deliberately does not claim that the rendered transition matches the user's intent - no runtime layer can decide that - but it confines that unbounded question to a single predicate over a finite object, which is what makes a preview a sufficient control rather than a formality. Seven safety invariants, derived from the fidelity condition rather than enumerated from experience, discharge the guarantee. Empirically, on a controlled N=60 adversarial suite the stack lifts pass rate by ~74 percentage points over a naive-ReAct baseline on two write-aggressive backing models, but by only ~3 points on a write-cautious one - evidence that single-model evaluations of agent safety stacks are close to unfalsifiable. The system is deployed; 108 production write operations across 8 chains back the failure taxonomy. Although the evaluation setting is public ledgers, the formalism and invariants apply to any probabilistic agent acting on irreversible external state.
Comments29 pages, 8 figures, 7 tables