一次 rollout 足矣:GUI 智能体的完全测试时适应
One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
浏览论文内容
中文总结 AI 辅助
针对GUI智能体在部署中无法获取真实标签且仅有一次尝试机会的约束,提出完全测试时适应方法SOLO,通过辅助模型筛选成功回合并重标注失败前缀,结合滑动窗口与自蒸馏更新适配器,在多个基准上提升成功率3-6个百分点。
中文摘要 AI 辅助
GUI 智能体以冻结权重部署,并丢弃其在工作中经历的一切。现有的更新智能体权重的方法假设了部署所不具备的条件:真实标签、超出单次尝试的 rollout(重试、采样、练习运行),或部署之外的学习阶段。由于 GUI 动作可能是不可逆的,部署的智能体在每次任务出现时只有一次尝试机会,按到达顺序处理,且每次尝试都至关重要。在任何时刻都无法获得真实标签。我们通过这些约束定义了 GUI 智能体的完全测试时适应,并将其与一种极简的权重空间方法 SOLO 配对。辅助模型读取每个回合:一个评判者选择其认为成功的回合,一个提议者-验证者对将失败回合的前缀重新标注为该前缀所完成的子任务。被接纳的回合进入一个短滑动窗口,每次接纳通过基于智能体自身预测的 top-K 自蒸馏更新一个小型适配器,前提是窗口中包含一个被评判为成功的回合。在由 WebArena、VisualWebArena 和 MobileWorld 构建的重复任务流上,SOLO 在 UI-TARS-7B 和 Qwen3-VL-8B 上相比冻结智能体将成功率提高了三到六个百分点,并在网络流上超过了两种设置内记忆方法。
英文摘要
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.
发表机构
- Concordia University(康考迪亚大学)
- Mila -- Québec AI Institute(Mila魁北克人工智能研究所)
- University of Toronto(多伦多大学)
- Shanghai University(上海大学)
机构由 AI 辅助整理,请以论文原文为准。