发表机构
Imperial College London; Indian Institute of Science Bangalore; Microsoft Security Response Center(帝国理工学院; 班加罗尔印度科学学院; 微软安全响应中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Planarian通过状态点抽象和快照、回滚、分支三种原语,管理智能体本地与远程状态,支持错误恢复与并行探索,任务质量提升高达15倍,开销仅3%。
AI 中文摘要
LLM智能体通过迭代地修改文件、调用本地工具以及与远程服务交互来解决复杂任务,这会在其本地环境和远程服务中修改状态。如今,智能体和用户必须显式地管理这些更改,无论是回退探索性操作还是从错误操作中恢复。安全地执行此操作需要协调一致的动作,然而当前的智能体框架缺乏统一的抽象和机制来一致且高效地管理本地和远程状态。我们描述了Planarian,一个具有状态管理的智能体运行时,它使智能体和用户能够从错误操作中恢复,并在一致的本地和远程环境状态上探索替代执行方案。Planarian引入了智能体状态点的抽象,即环境状态的一致、可恢复的时间点版本。Planarian向智能体和用户暴露了三种状态管理原语:(i)快照创建跨越本地和远程状态的新状态点,无需外部服务支持检查点:它依赖于高效的增量进程和文件系统快照来捕获本地沙盒状态,并透明地记录补偿操作以撤销远程状态更改;(ii)回滚通过恢复到先前的本地检查点并重放远程状态更改的补偿操作,将环境恢复到先前的状态点;(iii)分支从状态点创建多个隔离的分支,使智能体能够并行探索替代方案。我们展示了Planarian使智能体能够撤销错误并并行探索替代方案,将任务质量提升高达15倍,并允许用户仅以3%的开销从错误操作中恢复。
英文摘要
LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from erroneous ones. Doing so safely requires coordinated actions, yet current agent harnesses lack unified abstractions and mechanisms for managing local and remote state consistently and efficiently. We describe Planarian, an agent runtime with state management that enables agents and users to recover from erroneous actions and explore alternative executions over consistent local and remote environment state. Planarian introduces the abstraction of agent statepoints, which are consistent, restorable point-in-time versions of the environment state. Planarian exposes three state-management primitives to agents and users: (i) snapshot creates a new statepoint spanning local and remote state without requiring external services to support checkpoints: it relies on efficient incremental process and file system snapshotting to capture local sandboxed state, and transparently records compensating actions to undo remote state changes; (ii) rollback restores the environment to a previous statepoint by reverting to a prior local checkpoint and replaying compensating actions for remote state changes; and (iii) fork creates multiple isolated branches from a statepoint, enabling the agent to explore alternatives in parallel. We show that Planarian enables agents to undo mistakes and explore alternatives in parallel, improving task quality by up to 15x, and allows users to recover from erroneous actions with only 3% overhead.