SPIEval:评估大型语言模型作为处理分散个人信息的移动助手的能力
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
浏览论文内容
中文总结 AI 辅助
本研究推出基于五种认知能力的人工策划基准SPIEval,评估9种LLMs作为移动助手处理分散个人信息的能力,发现其准确率最高仅57.3%,存在信息定位等显著局限。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被部署为移动助手,其核心挑战是利用分散在多个应用程序(apps)中的个人信息来完成用户指令。然而,由于缺乏专用基准,它们的能力仍未得到充分了解。为解决这一差距,我们引入了SPIEval,这是一个由人工策划的基准,基于五种认知能力(即推理、消歧、整合、偏好推断和多意图分解)。SPIEval包含250个任务,分布在10个应用程序中的4335条个人记录,并通过21种工具支持多轮交互。分析表明,该基准具有多样化的场景、具有挑战性的任务、分散的信息、可控的环境和可验证的结果。我们评估了9种代表性LLMs,发现其仍有很大的改进空间。表现最佳的模型GPT-5.5(xhigh)仅达到57.3%的准确率,而最弱的仅达到16.4%。进一步分析显示,79%的失败源于不准确的信息定位,因为LLMs往往会承诺看似合理但不正确的信息,而非继续检索进行验证。我们还发现,不到2%的检索操作采用了高级搜索方法,并观察到不同模型之间搜索效率存在显著差异。这些发现揭示了当前基于LLM的移动助手的基本局限性,并推动了该方向的未来研究。数据和代码可在this https URL获取。
英文摘要
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.
发表机构
- Fudan University(复旦大学)
- Tencent Hunyuan Team(腾讯混元团队)
机构由 AI 辅助整理,请以论文原文为准。