发表机构
Salesforce AI Research(Salesforce AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对计算机使用代理,提出基于程序状态的StateAct多代理框架,主要代理用代码处理状态,GUI子代理辅助,支持验证,在OSWorld 2.0上提升了Claude Opus 4.8的成功率,且成本降低,实现从感知到推理的瓶颈转移。
AI 中文摘要
计算机使用代理通常通过增强感知来改进,即更好地读取屏幕截图并选择点击位置的模型。然而,屏幕截图只是底层程序状态的有损呈现。StateAct是围绕这一区别构建的代码优先的多代理框架。其主要代理通过代码直接处理程序状态,一个专用的GUI子代理处理少数需要的子目标的屏幕截图和点击交互。对程序状态的直接访问也支持验证。在OSWorld 2.0上,StateAct将Claude Opus 4.8的二进制成功率从20.6%提高到26.9%,部分成功率从54.8%提高到61.6%,每个任务的成本比仅由屏幕截图驱动的相同模型低约9倍。一般来说,将动作、验证和内存基于状态,即状态基础,将主要瓶颈从感知转向推理。
英文摘要
Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with program state by using code, while a dedicated GUI subagent handles screenshot-and-click interaction on the few subgoals that need it, just 28 of 108 tasks and 1.1% of main-agent steps. The same direct access to program state also supports verification: an independent finish gate double-checks the saved result for structural failures, e.g., output that is missing, unsaved, or written to the wrong path. To stay on track over hundreds of steps, the main agent hands subgoals to fresh subagents, keeping its own context focused. On OSWorld 2.0, StateAct lifts Claude Opus 4.8 from 20.6% to 26.9% on binary success, and from 54.8% to 61.6% on partial success, at ~ 9x lower cost per task than the same model driven by screenshots alone; a code-only variant with no GUI subagent reaches only 45.9% partial, below that screenshot-based baseline's 54.8%. In general, grounding action, verification, and memory in state, what we call state-grounding, shifts the main bottleneck from perception toward reasoning: failures depend more on what the agent thinks than on what it sees.