发表机构
Scale AI; Johns Hopkins University; University of Texas at Dallas(Scale AI; 约翰斯·霍普金斯大学; 德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体计算机使用中虚拟机状态无法合并的瓶颈,提出Spine-Branch协调框架,在200项长视距任务上提升成功率6.0%-16.5%、降低成本34%-70%,实现高效扩展。
AI 中文摘要
计算机使用智能体(Computer Use Agents, CUAs)正日益以多智能体系统形式部署,这类系统会将任务分解为多个子任务,由并行虚拟机(Virtual Machines, VMs)执行。然而,一个关键的物理瓶颈在于,两台虚拟机的状态无法合并。现有系统对此采用临时处理方式,未将其作为核心问题对待。本文提出面向多智能体计算机使用的脊柱-分支协调(Spine-Branch Coordination)框架,该框架将任务分解为“脊柱-分支”图结构:脊柱承载包含连续虚拟机状态的主任务流,分支任务并行执行以收集脊柱完成任务所需的信息;分支虚拟机在任务完成后即被丢弃,因此无需合并虚拟机状态。实验结果显示,在来自Odysseys的200项长视距任务上,针对三种CUAs主干模型,Spine-Branch框架较基线系统将成功率提升了6.0%至16.5%,同时将单任务成本降低了34%至70%,表明显式建模虚拟机状态合并约束可让多智能体计算机使用实现高效扩展。
英文摘要
Computer use agents (CUAs) are increasingly deployed as multi-agent systems that decompose a task into multiple subtasks executed across parallel virtual machines (VMs). However, a critical physical bottleneck is that the state of two VMs cannot be merged. Previous systems handle this ad-hoc rather than treating it as a first-class concern. We propose Spine-Branch Coordination for multi-agent computer use, a framework that decomposes a task into a "spine-branch" graph, where the spine carries the main task flow with continuous VM state and branch tasks execute in parallel to collect information the spine needs to complete the task. Branch VMs are discarded once their tasks finish, so no VM merging ever occurs. Experiments show that on 200 long-horizon tasks from Odysseys and across three CUA backbones, Spine-Branch improves success rate over the baseline system by 6.0% to 16.5%, while reducing per-task cost by 34% to 70%, indicating that explicitly modeling VM-state merging constraint enables multi-agent computer use to scale efficiently.