arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向时间干预下个人语言模型代理的用户条件评估

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

Pin Qian, Su Wang, Yihang Chen, Qiaolin Yu, Xiaoyuan Wang, Zhitong Guo, Zhicheng Wang, Junxian You

arXiv 2607.21635首次发表:更新:

发表机构

Carnegie Mellon University; Georgia Institute of Technology; Cornell University; University of Glasgow(卡内基梅隆大学; 佐治亚理工学院; 康奈尔大学; 格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究个人语言模型代理在时间干预下的用户条件评估,提出需在不同用户条件状态下重放时间干预并测量故障传播的协议,形式化四个条件,审查发现无满足条件的公开协议,进而提出最小基准设计和候选报告指标。

AI 中文摘要

个人代理维护着随每个用户而演变的记忆、学习技能、工具配置和策略状态。现有代理基准测试往往孤立地评估这些能力。我们认为个人代理评估需要不同的协议:在不同的持久用户条件状态下重放相同的时间干预,并测量故障如何在代理组件间传播。我们将此要求形式化为四个条件。对公开基准测试协议的审查发现没有一个满足所有条件。本文提出了最小基准设计和候选报告指标。

英文摘要

Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and safety benchmarks test static policy compliance. We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned states and measuring how failures propagate across agent components. We formalize this requirement as four conditions: explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, and variation in user-conditioned state. A focused audit of public benchmark protocols selected by explicit inclusion criteria identifies several close cases. Under our explicitly narrow operationalization, we did not find a protocol in that audited set satisfying all four conditions. This claim is scoped as a focused gap analysis with bounded literature coverage. This position paper proposes a minimal benchmark design and candidate reporting metrics for user-conditioned adaptation. The result is a concrete design requirement for future personal-agent evaluation, with metrics used as reporting tools for that requirement.

Comments9 pages, 2 figures, and 8 tables. Accepted for oral presentation at the ACM SIGKDD KDD 2026 Workshop on Personal Intelligence in the Agentic AI Era (PILA 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑