全量召回的代价是什么?智能体记忆系统的服务成本基准测试
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems
浏览论文内容
中文总结 AI 辅助
本研究基准测试 Mem0 等三种智能体记忆系统的服务成本,对比固定窗口与全量重发策略,发现其成本受内部行为驱动,且无系统同时在成本与准确率上占优。
中文摘要 AI 辅助
长期运行的对话智能体越来越依赖记忆系统,以避免每一轮都重发整个对话内容,但这类系统的服务成本尚未得到系统的基准测试。本研究将三种记忆系统(Mem0、Hindsight 与 Mastra Observational Memory)与两种参考策略——固定大小滑动窗口和重发完整对话文本——进行对比,涉及两种主干模型和最多 400 轮的对话,每一项成本测量都对应 665 个 LoCoMo 问题上的答案准确率。研究发现:第一,记忆系统的服务成本无法仅通过对话长度和消息大小预测,跟踪两种参考策略的回归模型对记忆系统的预测误差达 18%-69%,其成本由内部记忆行为驱动;第二,盈亏平衡分析显示,记忆系统是否以及何时比重发完整文本更具成本优势,高度依赖具体系统和主干模型,最便宜的系统在数十轮后即可实现,最昂贵的系统在 400 轮内始终无法实现;第三,没有系统能在成本和准确率两个维度同时占优,准确率范围为 21%-54%,主干模型对成本的影响程度与记忆系统相当。
英文摘要
Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory systems (Mem0, Hindsight, and Mastra Observational Memory) against two reference strategies -- a fixed-size rolling window and resubmitting the full transcript -- across two backbones and conversations of up to 400 turns, pairing every cost measurement with answer accuracy on 665 LoCoMo questions. First, a memory system's serving cost cannot be predicted from conversation length and message size alone: a regression that tracks the two reference strategies closely misses the memory systems by 18-69%, their cost driven instead by internal memory behavior. Second, a break-even analysis shows that whether -- and when -- a memory system becomes cheaper to serve than the full transcript is highly sensitive to the system and the backbone, from the first tens of turns for the cheapest to never within 400 turns for the most expensive. Third, no system wins on both axes: accuracy spans 21-54%, and the backbone choice drives cost as much as the memory system does.
发表机构
- Bricks Technology(布里克科技)
机构由 AI 辅助整理,请以论文原文为准。