arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DyadMem:智能体如何与用户协作的长期记忆基准

DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

Yifei Tao, Xinyu Zhong, Henry Hengyuan Zhao, Fanyi Wang, Tengda Guo, Wentao Qiu, Ying Wang, Liujian Tang

arXiv 2610.03020首次发表:更新:

发表机构

Nanyang Technological University; National University of Singapore; University of Hong Kong; Stepfun(南洋理工大学; 新加坡国立大学; 香港大学; Stepfun)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DyadMem提出用户条件关系智能体记忆(URAM)基准,包含大量多会话标注,揭示前沿模型在完整流水线问答中性能骤降,并验证URAM对20个模型均有积极效果。

AI 中文摘要

长期智能体不仅需要记住关于用户的真实信息,还需要记住随着共享历史的发展,特定智能体应如何与该用户协作。现有基准主要监督用户事实和偏好,或跨用户可复用的经验,使得这种关系特定的智能体记忆隐含化。此外,大多数先前工作仅通过长交互历史上的最终答案问答来评估模型,使得评估仍然不完整且不可靠。为此,我们引入了DyadMem,并提出了新定义的用户条件关系智能体记忆(URAM)。DyadMem沿相同的多会话轨迹联合标注用户侧记忆和URAM,产生6个记忆类别。总体而言,它包含3,065个情节、50,961个会话和61,210个问答实例,并具有广泛的会话级捕获和更新黄金标注、查询级召回支持,以及两种问答设置:黄金记忆和全流水线。在16个开放权重和4个专有模型中,黄金记忆问答始终表现强劲,而全流水线问答则急剧下降。这种差距明确支持了我们细粒度的评估设计。此外,多项定量结果进一步揭示了即使是前沿大语言模型也存在的低捕获召回率、不完整召回和不安全删除问题。我们进一步进行了严格的实验来验证URAM的有效性,并观察到对所有20个模型的积极影响。总之,DyadMem是一个双领域、全流水线的记忆基准,具有广泛的标注工作,以推动该领域的发展。

英文摘要

Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific agent memory implicit. Additionally, most prior works measure the model solely with final-answer QA over long interaction histories, making the assessment still incomplete and unreliable. To this end, we introduce DyadMem with the proposed new definition User-conditioned Relational Agent Memory (URAM). DyadMem jointly annotates user-side memory and URAM along the same multi-session trajectories, resulting in 6 memory categories. To summarize, it includes 3,065 episodes, 50,961 sessions, and 61,210 QA instances, with extensive session-level Capture and Update gold annotations, query-level Recall support, and two QA settings: Gold-Memory and Full-Pipeline. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong, yet Full-Pipeline QA drops sharply. Such a gap explicitly supports our fine-grained evaluation design. Additionally, several quantitative results further reveal low Capture recall, incomplete Recall, and unsafe-deletion issues arising from even the frontier LLMs. We further conduct a rigorous experiment to validate the effectiveness of our URAM and observe the positive effects for all 20 models. In summary, DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑