arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23035cs.AI

MobilePA-Bench:面向复杂真实世界任务的移动规划智能体基准测试

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao, Long Chen, Weigao Sun, Feida Zhu, Yiran Zhong, Steven Hoi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出 MobilePA-Bench 基准,针对现有移动智能体评估基准的盲区,从多维度测试移动规划智能体,发现前沿 LLM 在移动场景中存在可靠性问题,该基准可用于诊断与加速可靠移动智能体开发。

中文摘要 AI 辅助

随着端侧大语言模型(LLM)智能体发展为个人 copilots,移动操作系统已成为该范式的关键测试平台,因此严格的能力评估至关重要。然而现有基准分为两类,各有关键盲区:以图形用户界面(GUI)为中心的基准仅测试表层屏幕操作,却忽略后台工具使用与长程规划;而静态函数调用基准依赖离线API匹配,与真实运行时约束脱节。为弥合这一差距,我们提出 MobilePA-Bench——一种交互式、有状态、以工具为中心的基准,用于评估移动规划智能体的工具调用与规划能力。MobilePA-Bench 在可执行沙箱上运行,该沙箱维护实时应用数据库并返回结构化反馈,覆盖 13 个功能领域与 212 种真实移动工具。除基础工具使用外,它沿三个高级维度评估核心规划智能体:(1)子智能体协作——分解复杂任务并将专业工作委派给能干的子智能体;(2)记忆使用——调用存储的记忆、用户资料与过往偏好以解决隐式请求;(3)技能使用——调用预打包的复合技能而非从零规划每一步。大量实验表明,当前前沿 LLM 在移动场景中仍不可靠:在严格的工具顺序、权限限制与意外运行时错误下,性能急剧下降。通过将交互式函数调用沙箱与基于证据的验证相结合,MobilePA-Bench 既是实用的诊断基准,也是智能体强化学习的交互式基础,可加速可靠移动智能体的开发。

英文摘要

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present \textbf{MobilePA-Bench}, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning $13$ functional domains and $212$ realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: \emph{(1)~Sub-agent Collaboration}---decomposing a complex task and delegating specialized work to capable sub-agents; \emph{(2)~Memory Usage}---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and \emph{(3)~Skill Usage}---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.

发表机构

  • Alibaba(阿里巴巴)

机构由 AI 辅助整理,请以论文原文为准。

↑