arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PastForward:通过计算经验重用实现更快的端上GUI智能体

PastForward: Faster On-Device GUI Agents via Computational Experience Reuse

Taehwan Park, Changmin Lee, Hayeon Lee, Taesik Gong

arXiv 2609.32166首次发表:更新:

AI 中文总结

PastForward通过细粒度重用计算经验加速端上GUI智能体,在单次VLM前向验证多令牌提议并跨步骤复用KV状态,实现1.63-2.36倍延迟加速且保持任务成功率。

AI 中文摘要

在边缘设备上运行GUI智能体可以将敏感的屏幕和交互历史保留在本地,但每一步推理的计算成本使得部署面临挑战。现有的GUI智能体系统要么在每一步执行完整的视觉语言模型(VLM)推理,要么重用与先前任务匹配的粗粒度知识。然而,动态的移动环境和用户任务使得在没有额外微调或特定于任务的离线探索的情况下,难以充分利用先前的任务执行。为了应对这一挑战,我们提出了PastForward,一个通过验证的、细粒度的计算经验重用(这些经验在普通任务执行过程中积累)来加速GUI智能体的系统。在解码过程中,PastForward检索先前的输出序列作为设备自适应的多令牌提议,并在单次VLM前向传递中验证它们。在跨动作步骤中,它利用先前的GUI转换在设备执行当前动作时开始下一步推理,仅当预测屏幕与观察屏幕匹配时保留早期计算,并携带可重用的KV状态向前传递。我们在源自真实移动使用模式的AndroidWork负载上,使用多种VLM骨干网络在服务器和边缘平台上评估了PastForward。在设备上,PastForward实现了1.63-2.36倍的动作步骤延迟加速,同时保持了任务成功率。

英文摘要

Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks. However, dynamic mobile environments and user tasks make it difficult to fully utilize prior task executions without additional fine-tuning or task-specific offline exploration. To address this challenge, we present PastForward, a system that accelerates GUI agents through validated, fine-grained reuse of computational experience accumulated during ordinary task execution. During decoding, PastForward retrieves prior output sequences as device-adaptive multi-token proposals and verifies them in a single VLM forward pass. Across action steps, it uses prior GUI transitions to begin next-step inference while the device executes the current action, retains the early computation only when the predicted screen matches the observed screen, and carries reusable KV states forward. We evaluate PastForward on AndroidWorld workloads derived from real mobile usage patterns using multiple VLM backbones across server and edge platforms. On device, PastForward achieves action-step latency speedups of 1.63-2.36$\times$ while maintaining task success rates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑