发表机构
Georgia Institute of Technology; Stanford University; Simular(佐治亚理工学院; 斯坦福大学; Simular)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对重复性计算机任务,提出神经符号策略迭代方法,学习可重用策略,将稳定决策编码为代码,动态决策交由神经模型,显著提升可靠性和效率。
AI 中文摘要
许多计算机任务会重复出现:相同的工作流程会运行多次,每次都有新的输入和不同的起始状态。当前的计算机使用代理在每次运行中都会重新规划每一步,这使得它们在处理此类任务时成本高昂且不可靠。我们引入了神经符号计算机使用,其中重复出现的工作流程由学习到的策略执行,而不是由代理在每次运行时重新推导。该策略将跨运行稳定的决策(顺序、变量、循环和分支)固定在可执行代码中,并将依赖于观测的决策(如基础定位和状态检查)委托给神经模型。我们通过神经符号策略迭代来学习这些策略:从一条代理轨迹开始,执行策略,使用任务完成度和步骤级评判器诊断失败,并根据代理从失败点继续的信息,由编码模型修改代码,而无需访问基准评估器。对生成的参数和初始状态变体进行迭代使策略可重用,部署时的动作前验证器会保护每个改变状态的操作。在OSWorld-Verified和ScienceBoard上,学习到的策略在所有四个设置中取得了所有方法中最高的Pass^3,比基础代理高3.6-15.8个百分点,同时将每次运行成本降低了15-217倍,延迟降低了3.4-5.1倍。在OSWorld-Verified上,仅基于变体构建的策略能够迁移到保留的原始任务,在Pass^3上超过AutoRPA 8.6-17.5个百分点。
英文摘要
Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural models. We learn these policies with neuro-symbolic policy iteration: starting from one agent trajectory, it executes the policy, diagnoses failures with task-completion and step-level judges, and revises the code with a coding model informed by an agent's continuation from the point of failure, without access to the benchmark evaluator. Iterating on generated parameter and initial-state variants makes the policy reusable, and a pre-action verifier guards each state-mutating step at deployment. On OSWorld-Verified and ScienceBoard, the learned policies achieve the highest Pass^3 of all methods in all four settings, 3.6-15.8 points above the base agent, while cutting per-run cost by 15-217$\times$ and latency by 3.4-5.1$\times$. On OSWorld-Verified, policies built only on variants transfer to the held-out original tasks, exceeding AutoRPA by 8.6-17.5 points in Pass^3.
Comments26 pages, 8 figures, 10 tables