MemArena:面向移动端智能体个人记忆助手的大规模自我中心基准测试
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
浏览论文内容
中文总结 AI 辅助
本研究针对现有记忆基准测试的不足,构建了MemArena基准,评估了不同记忆后端对移动端个人记忆助手的影响,发现后端选择对内容准确性影响更大,权限感知访问失效,搜索延迟仅在阅读器规模极小时有影响。
中文摘要 AI 辅助
边缘部署的个人记忆助手必须使用开源权重模型在设备本地处理私密的人际对话。然而,现有的记忆基准测试往往未能充分测试密集活动交互、自我中心视角以及连贯多会话场景的组合情况。MemArena通过其MASim智能体模拟器构建了一个单场景对话基准测试来填补这些空白,该基准测试包含50个智能体,持续15天,共产生1030万条对话文本token,每个智能体每天产生24100条仅文本的自我观察token。借助交互历史,它在回忆、推理和可信赖性这六个评估维度上共同生成了真实值。我们评估了五个开源权重阅读器,分别使用Vanilla上下文、BM25-RAG、Oracle检索、Memobase和MemSearch作为记忆后端。有三个结果尤为突出:(1)记忆后端的选择对内容准确性影响更大:在Qwen3-0.6B模型上,从Memobase切换到MemSearch时,准确率提升了32.5/19.2个百分点,超过了MemSearch阅读器扩展带来的提升(10.6/6.8个百分点);(2)权限感知访问全面失效,Oracle后端泄露情况严重,其他后端则过于谨慎而不敢披露信息;(3)搜索延迟仅在阅读器规模极小时才会产生影响:在Spark GB10边缘节点上,内存搜索会增加适度且固定的87/7/48毫秒延迟(分别对应BM25-RAG/Memobase/MemSearch),对于大多数阅读器-后端组合而言,这只占TTFT的一小部分。代码、MASim模拟器和MemArena-L基准测试将在论文接收后发布。
英文摘要
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text tokens, 24.1K text-only ego-observed tokens/agent/day). With the interaction history, it co-generates ground truth over six recall, reasoning, and trustworthiness evaluation dimensions. We evaluate five open-weight readers with Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch as memory backends. Three results stand out: (1) Memory-backend choice matters more for content accuracy: At Qwen3-8B, moving from Memobase to MemSearch gains +22.1/+21.2 pp, whereas scaling the reader to Qwen3-32B-AWQ gains at most +3.5/+4.4 pp under either backend. (2) Permission-aware access fails in two distinct modes: Oracle leaks heavily, while the other backends fail to surface the protected fact. (3) Search latency bites only at small readers: on a Spark GB10 edge node, memory-search adds a moderate and fixed 87/8/51 ms (BM25-RAG/Memobase/MemSearch) that composes a small part of TTFT for most reader-backend combinations. We release code, the MASim simulator, and the MemArena-L benchmark at https://github.com/dereksodo/MemArena-Bench.
发表机构
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。