arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Jev-Mobile:Jev作为移动GUI智能体的执行器

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Linghua Zhang

arXiv 2609.30186首次发表:更新:

AI 中文总结

Jev-Mobile通过低频VLM规划与高频轻量级执行解耦,在AndroidWorld上以79%成功率实现32.7%时间缩减和73.4%成本降低,兼顾效率与性能。

AI 中文摘要

视觉语言模型(VLMs)已成为自主移动GUI智能体的常见基础,但大多数现有系统在几乎每个交互步骤中都依赖VLM进行规划和动作落地,导致大量延迟和模型服务成本。我们提出Jev-Mobile,将这一范式转变为低频VLM规划和高频轻量级执行:VLM指定局部目标,可访问性树定义结构化的可执行动作空间,而Jev(一种快速类型化决策模型)在该空间内反复选择动作。这种设计允许在单个VLM决策下执行多个GUI动作,减少昂贵的VLM推理,同时保持自适应交互。在完整的AndroidWorld任务套件上,Jev-Mobile实现了79%的任务成功率,而SeeAct-V为78%,逐步VLM基线为84%。在成功轨迹中,相对于逐步VLM,Jev-Mobile将平均端到端执行时间减少了32.7%,平均模型API成本减少了73.4%。这些结果表明,将高层VLM推理与低层动作执行解耦,可以在保持竞争性任务性能的同时,显著提高移动GUI智能体的效率。

英文摘要

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑