arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从交互轨迹到持久技能:计算机使用智能体的在线演化

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

Longtao Hu, Xiao Liang, Linchao Zhu

arXiv 2609.04869首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Zhejiang University(电子科技大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出在线技能演化框架,构建可复用过程的持久化版本库,经实验验证其在 OSWorld 四域中较空库对照组得分更高,收益具条件性且重复修订无法保证恢复。

AI 中文摘要

计算机使用智能体可在图形界面中执行日益复杂的任务,但其交互体验通常是短暂的:从一次 rollout 中获取的过程性知识未被系统地保留、优化并在后续任务中复用。现有技能库提供外部过程性知识,但其相较于不使用技能的相同智能体的增量价值,以及在重复交互下的纵向动态仍未得到充分表征。我们提出一种在线技能演化框架,该框架将交互轨迹与评估器反馈转换为可复用过程的持久化、版本化库。每次迭代针对冻结的库快照执行,经证据引导的技能更新可在后续迭代中使用,且无需修改模型参数。我们在四个 OSWorld 应用领域中,于相同的固定动作生成与 GUI 接地栈、任务集及迭代周期下,将完整的演化库系统与配置匹配的空库对照组进行对比。经过五次迭代的空库预热后,完整系统在所有四个观测域运行中达到更高的预热后平均评估器得分,平均差异范围为 5.7 至 18.6 个百分点,且具有依赖于域的时间稳定性。在 GIMP 中,经溯源感知分析显示,存在跨源任务边界的检索与修订 churn,其中重复的已接受编辑无法恢复源任务。这些发现将演化技能库表征为可审计的共享过程记忆,可改进固定的计算机使用栈,同时表明其收益具有条件性,且重复修订无法保证恢复。代码已发布至 this https URL。

英文摘要

Computer-use agents can execute increasingly complex tasks in graphical interfaces, but their interaction experience is typically transient: procedural knowledge acquired from one rollout is not systematically retained, refined, and reused in later tasks. Existing skill libraries provide external procedural knowledge, yet their incremental value over the same agent operating without skills, as well as their longitudinal dynamics under repeated interaction, remain insufficiently characterized. We present an online skill-evolution framework that converts interaction trajectories and evaluator feedback into a persistent, versioned library of reusable procedures. Each iteration executes against a frozen library snapshot, and evidence-guided skill updates become available in subsequent iterations without changing model parameters. We compare the full evolving-library system with a configuration-matched empty-library control across four OSWorld application domains under the same fixed action-generation and GUI-grounding stack, task sets, and iteration horizons. Following a five-iteration empty-library warm-up, Full attains a higher post-warm-up mean evaluator score in all four observed domain runs, with mean differences ranging from 5.7 to 18.6 percentage points and domain-dependent temporal stability. In GIMP, provenance-aware analysis reveals retrieval across task-of-origin boundaries and revision churn, where repeated accepted edits fail to recover the originating task. These findings characterize evolving skill libraries as auditable, shared procedural memory that can improve a fixed computer-use stack, while showing that their benefits are conditional and repeated revision does not guarantee recovery. Code is released at https://github.com/LongtaoHu/Skill-Evo4GUI.

Comments10 pages, 2 figures, and 2 tables. Code: https://github.com/LongtaoHu/Skill-Evo4GUI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑