arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StateFork:面向智能体探索的可分支基础设施

StateFork: Branchable Infrastructure for Agent Exploration

Jiakai Xu, Tianle Zhou, Georgios Liargkovas, Danielle Gillai, Ruizhe Fu, Patrick Shen, Eugene Wu, Kostis Kaffes

arXiv 2609.38648首次发表:更新:

发表机构

Columbia University; Google(哥伦比亚大学; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对终端智能体探索中环境状态分支恢复的难题,提出StateFork控制平面与Waypoint检查点基底,实现观测等价且物理高效的分支,在Terminal-Bench上提升任务完成率并加速探索。

AI 中文摘要

AI智能体通过探索多条轨迹来提高任务成功率,但对于计算机使用型智能体,每条轨迹都会修改外部环境状态。仅当恢复在观测上等价(即未来动作产生相同观测)时,从中间点分支才是正确的;且仅当创建、恢复和丢弃分支状态在物理上高效时,分支才具有实用性。我们针对使用终端的智能体研究该问题,此类任务会修改文件、shell上下文、运行中的进程和本地服务。我们提出StateFork,一个逻辑控制平面,将探索策略与物理状态物化解耦,在多种执行基底上暴露会话、命令、快照、恢复和清理操作。我们还构建了Waypoint,一个用于终端执行会话的检查点/恢复基底,结合了文件系统分层、进程检查点以及持久的终端兼容命令会话。StateFork和Waypoint共同通过结合样本高效的搜索与对正确执行状态的高效恢复,提升了终端智能体的探索能力。在Terminal-Bench上,通过StateFork和Waypoint进行基于分支的探索,在相同访问节点预算下相比pass@20提升了任务完成率,且使用Waypoint相比其他执行基底实现了高出26%的任务准确率,同时探索速度最高提升70%。这些结果表明,观测等价且物理高效的执行会话是探索性AI智能体的关键系统抽象。

英文摘要

AI agents improve task success by exploring multiple trajectories, but for computer-use agents each trajectory modifies external environment state. Branching from an intermediate point is correct only when restoration is observation-equivalent - future actions produce the same observations - and practical only when creating, restoring, and discarding branch states is physically efficient. We study this problem for terminal-using agents, where tasks modify files, shell context, running processes, and local services. We introduce StateFork, a logical control plane that separates exploration policies from physical state materialization, exposing sessions, commands, snapshots, restores, and cleanup over multiple execution substrates. We also build Waypoint, a checkpoint/restore substrate for terminal execution sessions that combines filesystem layering, process checkpointing, and a persistent terminal-compatible command session. Together, StateFork and Waypoint improve terminal-agent exploration by combining sample-efficient search with efficient restoration of the right execution state. On Terminal-Bench, branch-based exploration through StateFork and Waypoint improves task completion over pass@20 at the same visited-node budget, and using Waypoint achieves 26% higher task accuracy than other execution substrates while completing exploration up to 70% faster. These results show that observation-equivalent, physically efficient execution sessions are a key systems abstraction for exploratory AI agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑