arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Setoka:面向异构数据下个性化智能体分层用户理解的基准测试集

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou

arXiv 2607.27056首次发表:更新:

AI 中文总结

本研究提出Setoka基准测试集,评估记忆增强型个性化智能体的分层用户理解能力,发现现有系统在需整合异构信息的任务中性能下降,推动相关记忆机制设计。

AI 中文摘要

个性化智能体正越来越多地被应用于辅助用户完成各类任务。有效的个性化辅助不仅需要从存储在智能体记忆中的过往交互中检索明确事实,还需要推断抽象的个人特征。然而,现有的记忆基准测试集主要评估智能体是否能检索对话历史中明确陈述的信息,无法对更深层次的用户理解提供有效评估。在本研究中,我们提出了Setoka,这是一个用于评估具备分层用户理解能力的记忆增强型个性化智能体的基准测试集,相关数据来自异构数据源。Setoka基于认知心理学和人格心理学的理论,定义了四个层次的用户理解,即语义记忆、情景记忆、行为模式和人格特质。此外,为实现真实且隐私保护的评估,我们设计了一个基于心理测量学的流水线,可大规模合成多样、连贯的异构用户数据及查询。最后,我们利用Setoka评估了3种语言模型结合5种记忆系统在10个合成用户上的表现。我们的综合评估显示,现有系统在语义记忆检索上表现良好,但在情景记忆任务上性能下降;在处理需要整合随时间分散的异构碎片化信息的行为模式和人格特质理解任务时,性能下降更为明显。这些发现表明,用户理解不能仅靠简单的事实检索来实现,这为设计用于跨源整合及长期用户行为抽象的记忆机制提供了动力。

英文摘要

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑