arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13900cs.DBcs.AIcs.CLcs.LG

智能体事务:面向ACID兼容的智能体系统

Agentic Transaction: Towards ACID-Compliant Agent Systems

Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ACID兼容的智能体事务框架,开发对应数据智能体,在基准测试中较含Claude Code的现有智能体提升10.6%,为构建可信可扩展AI智能体开辟新方向。

中文摘要 AI 辅助

大型语言模型(LLM)智能体正从对话助手演变为通过推理、工具使用、代码生成和工作空间操作执行长周期任务的自主系统。随着智能体在持久环境和多步骤工作流中运行,它们面临与事务数据库系统所解决的类似挑战:可靠执行、一致结果、安全并发和持久状态管理。我们引入智能体事务的概念,并提出一种ACID兼容的智能体系统框架,该框架通过四个语义保障重新解释经典ACID属性用于智能体执行:语义原子性、语义一致性、语义隔离性和语义持久性。这些属性共同为构建可靠的智能体系统提供了原则性基础,尽管存在模型不确定性和动态执行环境。为实例化该框架,我们开发了一种ACID兼容的数据智能体,它通过事务式探索-执行-验证循环、事务式技能中心、基于置信度分歧的验证、语义依赖感知的隔离以及感知事务的语义状态管理来实现这些保障。在广泛使用的基准上的实验结果表明,我们的系统比包括Claude Code在内的最先进智能体实现了10.6%的提升。这项工作为扩展事务原则和系统架构以构建可信、可扩展和自演进的AI智能体系统开辟了更广泛的研究议程。

英文摘要

Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. As agents increasingly operate over persistent environments and multi-step workflows, they face challenges analogous to those addressed by transactional database systems: reliable execution, consistent outcomes, safe concurrency, and durable state management. We introduce the concept of an agentic transaction and propose an ACID-compliant agent system framework that reinterprets the classical ACID properties for agent execution through four semantic guarantees: Semantic Atomicity, Semantic Consistency, Semantic Isolation, and Semantic Durability. Together, these properties provide a principled foundation for building reliable agent systems despite model uncertainty and dynamic execution environments. To instantiate this framework, we develop an ACID-compliant data agent that realizes these guarantees through transactional exploration-execution-validation cycles, transactional skill hubs, confidence divergence-based validation, semantic dependency-aware isolation, and transaction-aware semantic state management. Experimental results on widely used benchmarks show that our system achieves a 10.6% improvement over state-of-the-art agents, including Claude Code. This work opens a broader research agenda on extending transactional principles and system architectures toward building trustworthy, scalable, and self-evolving AI agent systems.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑