arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DolphinBench:绘制智能体记忆的帕累托前沿

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Soumil Rathi, Deshraj Yadav, Taranjeet Singh

arXiv 2609.24971首次发表:更新:

发表机构

Mem0(Mem0)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DolphinBench通过任务完成直接评估智能体记忆,包含三个角色各约50万词元历史,验证200个任务,并要求同时报告准确性、成本和延迟,以全面绘制记忆系统的帕累托前沿。

AI 中文摘要

当前的智能体常常需要基于长期记忆和随时间推移的上下文回忆来执行现实世界中的行动。然而,目前大多数记忆基准测试都是为对话式问答格式设计的,在这种格式中,问题本身即提示了必须检索某个事实,且往往还提示了具体是哪个事实。此外,基准测试很少要求提交结果除准确性之外的指标,这使得记忆系统能够通过不合理的时间/成本权衡来获得更高的分数。我们提出了DolphinBench,一个通过智能体任务完成情况直接评估记忆的基准测试。DolphinBench包含三个知识工作角色,每个角色约有50万词元的用户消息,并评估智能体在依赖于该历史信息的任务上的表现。我们通过运行带有和不带相关历史的智能体来验证每个角色的全部200个任务,要求带有历史时成功,不带历史时失败。最后,我们要求所有评估在报告准确性的同时报告总成本和延迟,这使我们能够全面评估智能体记忆系统。现有的记忆基准测试没有同时结合这三者。数据集和评估代码可在以下网址获取:此https URL。

英文摘要

Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.

Comments6 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑