arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeltaReplay:面向移动GUI智能体的任务相对记忆复用

DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents

Yudong Bai, Yihong Chen, Quanming Yao, Yaqing Wang

arXiv 2610.11707首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DeltaReplay是一种步骤级记忆复用框架,通过拆分动作、按页面一致性和动作通用性匹配轨迹,在AndroidWorld和SPA-Bench上分别将任务成功率提升最高10.3和25.0个百分点,解决了移动GUI智能体的记忆复用困境。

AI 中文摘要

带记忆增强的移动GUI智能体会存储成功的执行轨迹并在后续任务中复用,但存储的轨迹很少能与新任务完全匹配。新任务可能使用不同的参数、仅与存储轨迹共享部分步骤,或在记忆中无相关记录。强制智能体使用无关记忆会误导它,而丢弃仍可能有用的记忆则会使其失去过往经验的指导。为解决这一困境,我们提出DeltaReplay,一种无需修改现有记忆的步骤级记忆复用框架。我们发现,存储记录的可复用部分并非由记录本身决定,而是由其与新任务的关系决定,主要通过两个因素:页面级一致性和动作级通用性。因此,我们将执行轨迹存储为转移图中的路径,其节点(页面)和边(页面间的动作)捕获这两个因素。在复用时,每条边的动作会被拆分为与任务无关的操作和与任务相关的参数。DeltaReplay随后将每个记录步骤与新任务及当前屏幕进行比较,决定是遵循该步骤、替换其参数后执行,还是交由基础智能体处理。在AndroidWorld和SPA-Bench上,DeltaReplay使采用相同骨干网络的基础智能体的任务成功率分别提升了高达10.3和25.0个百分点。这些结果表明,在每一步决定如何使用检索到的记忆,能让智能体即使从部分匹配的轨迹中也能获益。

英文摘要

Memory-augmented mobile GUI agents store successful execution trajectories and reuse them in later tasks, but a stored trajectory rarely matches a new task exactly. The new task may use different parameters, share only some of its steps with a stored trajectory, or have no relevant record in memory. Forcing the agent to use irrelevant memory can mislead it, whereas discarding memory that may still be useful deprives it of guidance from past experience. To address this dilemma, we propose DeltaReplay, a step-level memory reuse framework that decides how to use existing memory without modifying it. We observe that the reusable part of a stored record is determined not by the record itself but by its relation to the new task, mainly through two factors: page-level consistency and action-level generality. We therefore store execution trajectories as paths in a transition graph, whose nodes (pages) and edges (actions between pages) capture these two factors. At reuse time, the action on each edge is split into a task-independent operation and task-specific parameters. DeltaReplay then compares each recorded step with the new task and the current screen, and decides whether to follow it, execute it after replacing its parameters, or leave it to the base agent. On AndroidWorld and SPA-Bench, DeltaReplay improves the task success rate over a base agent with the same backbone by up to 10.3 and 25.0 percentage points, respectively. These results indicate that deciding at each step how to use retrieved memory lets agents benefit even from partially matching trajectories.

Comments22 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑