arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.00269cs.AI

Mnemosyne: 用于验证和修复AI生成工作流的代理事务处理

Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows

Edward Y. Chang, Longling Geng

首次发表
浏览论文内容

中文总结 AI 辅助

提出代理事务处理(ATP)模型,将生成动作视为不可信提案,通过确定性约束集验证后才提交,并实现运行时修复,确保状态正确性独立于生成层。

中文摘要 AI 辅助

LLM、求解器和代理团队越来越多地生成工作流动作、修复和计划,但生成的动作可能在语法上有效,却可能过时、不可行、冲突或破坏触发修复的证据。我们引入代理事务处理(ATP),一种事务模型,它将生成的动作视为不可信提案,直到它们通过声明的、可执行的约束集C下的确定性准入。原则是双重的:提案不是真理,且没有提案能预见所有干扰:任何事物都可以提议,但只有运行时接受并提交,当不可预见的干扰发生时,它会在边界内反应性地修复,而不是信任新的提案。相对于C,提交状态正确性变得独立于提议层的胜任能力、诚实性或学习能力。我们在Mnemosyne中实现了ATP,这是一个运行时系统,具有仅追加的转换日志、有效状态投影、依赖安全补偿和活动提交记录,并证明了相对于C的四个安全属性(权限分离、串行等效生成准入、证据保留修复和义务包含),以及其局部修复协议(LCRP)的有界反应性修复保证。一个可复现的工件在九个伪造测试中拒绝了目标违规,同时仍然接受有效工作,投影和验证开销低于6%,并且有界局部修复比全局重计算编辑的操作数量少一个数量级。Mnemosyne是开源的:此 https URL。

英文摘要

LLMs increasingly generate workflow actions and repairs that may be well formed yet stale, infeasible, conflicting, or destructive of their own evidence. We introduce Agentic Transaction Processing (ATP), which treats generated actions as untrusted proposals until deterministic admission accepts them under an executable constraint set C. Its two-sided principle is: a proposal is not truth, and no proposal foresees every disruption. Anything may propose, but only the runtime admits and commits; unforeseen disruptions trigger bounded reactive repair whose output re-enters admission. Mnemosyne realizes ATP with an append-only transition log, effective-state projection, dependency-safe compensation, and active contract records. Under stated assumptions and relative to C, we prove four safety properties (authority separation, serial-equivalent generative admission, evidence-preserving repair, and obligation containment) and establish bounded reactive repair. Across nine safety benchmarks and a four-case Temporal SDK comparison, ATP rejects every targeted violation while admitting valid work. A matched data-size sweep measures 5.2-6.6% incremental throughput cost over the same local durable commit path and exposes local saturation. In a companion scheduling harness, local repair edits nearly an order of magnitude fewer operations than global recompute; a 12-scenario interruption stress test rejects every stale recovery candidate without losing observations or producing invalid commits. Two bounded pilots route 80 proposals from four heterogeneous LLMs through the same gate with zero invalid commits; 24 of 40 mid-execution proposals are admitted and 16 are rejected, including four explicit safety rejections.

发表机构

  • Stanford University(斯坦福大学)
  • QuadriumAI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑