发表机构
Carnegie Mellon University; Georgia Institute of Technology; Cornell University; University of Glasgow(卡内基梅隆大学; 佐治亚理工学院; 康奈尔大学; 格拉斯哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究个人语言模型代理在时间干预下的用户条件评估,提出需在不同用户条件状态下重放时间干预并测量故障传播的协议,形式化四个条件,审查发现无满足条件的公开协议,进而提出最小基准设计和候选报告指标。
AI 中文摘要
个人代理维护着随每个用户而演变的记忆、学习技能、工具配置和策略状态。现有代理基准测试往往孤立地评估这些能力。我们认为个人代理评估需要不同的协议:在不同的持久用户条件状态下重放相同的时间干预,并测量故障如何在代理组件间传播。我们将此要求形式化为四个条件。对公开基准测试协议的审查发现没有一个满足所有条件。本文提出了最小基准设计和候选报告指标。
英文摘要
Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and safety benchmarks test static policy compliance. We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned states and measuring how failures propagate across agent components. We formalize this requirement as four conditions: explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, and variation in user-conditioned state. A focused audit of public benchmark protocols selected by explicit inclusion criteria identifies several close cases. Under our explicitly narrow operationalization, we did not find a protocol in that audited set satisfying all four conditions. This claim is scoped as a focused gap analysis with bounded literature coverage. This position paper proposes a minimal benchmark design and candidate reporting metrics for user-conditioned adaptation. The result is a concrete design requirement for future personal-agent evaluation, with metrics used as reporting tools for that requirement.
Comments9 pages, 2 figures, and 8 tables. Accepted for oral presentation at the ACM SIGKDD KDD 2026 Workshop on Personal Intelligence in the Agentic AI Era (PILA 2026)