发表机构
The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PhoneCLI将应用GUI导航编译为可调用命令,离线构建语义地图生成确定性重放,在线以零VLM成本执行,失败回退VLM,在AndroidLab和AndroidWorld上提升成功率并减少步骤与令牌消耗。
AI 中文摘要
移动图形用户界面智能体通过感知-行动循环运行:在每一步中,它们截取设备屏幕,调用视觉语言模型(VLM),并发出一个动作。这一过程缓慢、昂贵且脆弱,然而其大部分工作只是导航——而日常导航是静态的、有序的且无限重复的。我们提出了PhoneCLI,它将应用的图形用户界面导航编译为可调用命令,无需任何应用内部API、运行时插桩或模型训练。离线时,PhoneCLI从外部探索目标应用,并将其屏幕、交互元素和导航边蒸馏为语义标注的地图;每个屏幕产生一个确定性命令:一个到达该屏幕的重放序列。在线时,智能体选择一个命令,在执行前验证它,然后在亚秒级时间内以零VLM成本确定性执行;开放式交互和编译路径的任何失败都回退到内嵌的VLM解释器,即纯VLM智能体,因此编译只会带来帮助。在AndroidLab上,PhoneCLI提高了任务成功率,同时减少了步骤和令牌消耗,并迁移到AndroidWorld的官方M3A智能体上,带来一致的效率提升。PhoneCLI编译的是应用的导航而非单次运行,因此它服务于新任务,而不仅仅是重复任务。
英文摘要
Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld's official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app's navigation rather than one run, so it serves new tasks, not only repeated ones.