MobilePA-Bench:面向复杂真实世界任务的移动规划智能体基准测试
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
浏览论文内容
中文总结 AI 辅助
该研究提出 MobilePA-Bench 基准,针对现有移动智能体评估基准的盲区,从多维度测试移动规划智能体,发现前沿 LLM 在移动场景中存在可靠性问题,该基准可用于诊断与加速可靠移动智能体开发。
中文摘要 AI 辅助
随着端侧大语言模型(LLM)智能体发展为个人 copilots,移动操作系统已成为该范式的关键测试平台,因此严格的能力评估至关重要。然而现有基准分为两类,各有关键盲区:以图形用户界面(GUI)为中心的基准仅测试表层屏幕操作,却忽略后台工具使用与长程规划;而静态函数调用基准依赖离线API匹配,与真实运行时约束脱节。为弥合这一差距,我们提出 MobilePA-Bench——一种交互式、有状态、以工具为中心的基准,用于评估移动规划智能体的工具调用与规划能力。MobilePA-Bench 在可执行沙箱上运行,该沙箱维护实时应用数据库并返回结构化反馈,覆盖 13 个功能领域与 212 种真实移动工具。除基础工具使用外,它沿三个高级维度评估核心规划智能体:(1)子智能体协作——分解复杂任务并将专业工作委派给能干的子智能体;(2)记忆使用——调用存储的记忆、用户资料与过往偏好以解决隐式请求;(3)技能使用——调用预打包的复合技能而非从零规划每一步。大量实验表明,当前前沿 LLM 在移动场景中仍不可靠:在严格的工具顺序、权限限制与意外运行时错误下,性能急剧下降。通过将交互式函数调用沙箱与基于证据的验证相结合,MobilePA-Bench 既是实用的诊断基准,也是智能体强化学习的交互式基础,可加速可靠移动智能体的开发。
英文摘要
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present \textbf{MobilePA-Bench}, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning $13$ functional domains and $212$ realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: \emph{(1)~Sub-agent Collaboration}---decomposing a complex task and delegating specialized work to capable sub-agents; \emph{(2)~Memory Usage}---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and \emph{(3)~Skill Usage}---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
发表机构
- Alibaba(阿里巴巴)
机构由 AI 辅助整理,请以论文原文为准。