PAUSE:统一服务环境中面向用户的个人AI助手基准测试
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
浏览论文内容
中文总结 AI 辅助
研究人员针对现有个人AI助手基准测试的缺陷,提出PAUSE基准测试,采用多机制评估框架,实验发现现有先进模型在相关场景任务完成率不足70%,还提出可扩展生成任务的合成流水线。
中文摘要 AI 辅助
个人AI助手正日益作为面向任务、工具增强的智能体被部署于统一服务环境中,以支持用户日常活动。在实际场景中,这类助手必须对持续的用户状态进行推理,尊重用户特定的配置与权限,并在多个服务间维持长周期、感知约束的交互。然而现有基准测试常割裂服务上下文或抽象用户状态,限制了其在实际服务场景中评估以用户为中心的个人助手行为的能力。我们提出PAUSE,一个用于评估有状态、服务集成环境中个人AI助手的以用户为中心的基准测试。PAUSE通过要求智能体在异构用户拥有的资源间协调行动,同时保持与环境状态、多轮交互的授权约束的一致性,捕捉了现实助手部署的核心挑战。该基准测试通过逼真的用户模拟纳入显式的用户-智能体交互,支持超越静态工具执行的评估。为支持原则性与可复现的评估,PAUSE采用与任务特征对齐的多机制评估框架:开放式服务管理任务使用语义和轨迹级行为指标评估,而约束密集型任务则采用确定性、基于状态的验证。基准测试结果显示,即便是最先进的专有模型在需要有状态推理和配置感知的场景中,任务完成率也未达到70%,揭示了一致且可解释的失败模式。最后,我们提出一个以用户为中心的合成流水线,可规模化生成连贯的服务环境、用户配置和可靠标注的任务,支持基准测试的可扩展性与未来研究。
英文摘要
Personal AI assistants are increasingly deployed as task-oriented, tool-augmented agents that operate within unified service environments to support everyday user activities. In realistic settings, such assistants must reason over persistent user state, respect user-specific configurations and permissions, and sustain long-horizon, constraint-aware interactions across multiple services. Existing benchmarks, however, often fragment service contexts or abstract away user state, limiting their ability to evaluate user-centric personal assistant behavior in realistic service settings. We introduce PAUSE, a user-centric benchmark for evaluating personal AI assistants in stateful, service-integrated environments. PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state, authorization constraints over multi-turn interactions. The benchmark incorporates explicit user-agent interaction via realistic user simulation, enabling evaluation beyond static tool execution. To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification. Benchmark results show that even state-of-the-art proprietary models fail to reach 70% task completion on scenarios requiring stateful reasoning and configuration awareness, revealing consistent and interpretable failure patterns. Finally, we present a user-centric synthesis pipeline that enables scalable generation of coherent service environments, user configurations, and reliably annotated tasks, supporting benchmark extensibility and future research.
发表机构
- University of Alberta(阿尔伯塔大学)
机构由 AI 辅助整理,请以论文原文为准。