何时选择取代提取?基于类型化决策模型的智能体记忆预注册测试
When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
浏览论文内容
中文总结 AI 辅助
本研究通过预注册实验证明,在紧凑预算下,类型化决策模型Jev选择原始对话轮次可替代LLM提取记忆,且成本更低,并解释了不同预算下重排序收益差异的原因。
中文摘要 AI 辅助
对话记忆是否需要大语言模型提取的事实,还是选择正确的原始轮次就足够了?已发表的研究结果存在分歧。基于提取的系统报告了从蒸馏事实中获得的收益。近期研究发现,具有良好排序的原始历史同样有效,但关于排序是否重要存在分歧。我们在保留的LoCoMo对话和LongMemEval上进行了预注册研究。在LoCoMo的紧凑预算下,通过单次调用类型化决策模型Jev选择的原始轮次不劣于大语言模型提取记忆(单侧95%界限为-3.0点,对比-5点边际)。盲人人工评分缩小了边际但未改变结果。原始轮次的写入成本低3,061倍,且结果在第二个答案模型上仍然成立。在本研究中,随着预算增加,重排序的收益缩小。当保留30个候选中的3个时,它在LoCoMo上增加17.4点,在LongMemEval上增加9.1点。在宽松预算下,它增加1.5和1.1点,且提取系统更准确。这解释了已发表结果为何分歧。在匹配上下文下,Jev的选择准确性与大语言模型重排序器相当(非劣性界限-2.0),延迟仅为三分之一,且比多调用图遍历更准确。重排序降低了正确的弃权(不执行)。计划、代码和评分答案已发布。
英文摘要
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.