超越内存排行榜:将科学记忆评估为预算上下文恢复
Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
浏览论文内容
中文总结 AI 辅助
研究提出两个全文科学记忆基准测试PAIM和PTr,评估八个内存/检索系统,发现内存排行榜受多种因素影响,如摄取粒度等。结果显示不同系统表现因条件而异,还表明大语言模型评判排名与人工一致,应将科学记忆评估为预算上下文恢复并开源相关资源。
中文摘要 AI 辅助
长期记忆正成为大语言模型智能体的核心组件,但大多数内存基准测试评估的是对话或简短摘要,而研究智能体需要从完整的科学论文中恢复证据。我们引入了两个全文科学记忆基准测试,即公共人工智能记忆(PAIM;81篇论文,66个问题)和公共变压器(PTr;252篇论文,98个问题)。我们评估了八个内存/检索系统,包括我们自己提出的Theoria,以及一个无检索基线。我们的结果表明,如果没有完整的协议,内存排行榜是无法解释的:摄取粒度、原始文本保留、检索预算、检索模态、评分审核和评判选择都会影响结果。例如,在PAIM上,Graphiti令人信服地获胜,但每个查询使用260万个字符的检索上下文,在控制检索预算后,领先优势消失。在PTr上,对于可以干净地添加BM25检索的系统,稀疏-密集混合是最显著的单一干预措施:Simple RAG、Mem0和Theoria的混合变体并列领先,相差0.03分。多评判和人工并排校准表明,作为评判的大语言模型排名在前沿评判中是一致的,并且与人工评估一致,在十分制上的有效分辨率约为一分。我们认为,科学记忆应该被评估为有预算的、模态感知的上下文恢复,而不是无约束的架构排行榜,并且我们发布了数据集、工具、原始输出、评判和脚本,以重现我们的结果并作为此类评估的工具。我们的代码可在此http URL上获取,数据集可在此http URL上获取。
英文摘要
Long-term memory is becoming a core component of LLM agents, but most memory benchmarks evaluate conversations or compact summaries, while research agents need to restore evidence from full scientific papers. We introduce two full-text scientific-memory benchmarks, Public AI Memory (PAIM; 81 papers, 66 questions) and Public Transformers (PTr; 252 papers, 98 questions). We evaluate eight memory/retrieval systems, including our own proposed Theoria, plus a no-retrieval baseline. Our results show that memory leaderboards are not interpretable without the full protocol: ingestion granularity, raw-text preservation, retrieval budget, retrieval modality, rubric audit, and judge choice all affect the outcome. For example, on PAIM Graphiti wins convincingly but uses 2.6M characters of retrieved context per query, and after controlling for retrieval budget the lead disappears. On PTr, for the systems where BM25 retrieval can be added cleanly, the sparse-dense hybrid is the single most significant intervention: hybrid variants of Simple RAG, Mem0, and Theoria tie for the lead within 0.03 points. Multi-judge and human side-by-side calibration show that LLM-as-a-judge rankings are consistent across frontier judges and agree with human evaluation, with an effective resolution of roughly one point on a ten-point scale. We argue that scientific memory should be evaluated as budgeted, modality-aware context restoration rather than as an unconstrained architecture leaderboard, and we release the datasets, harness, raw outputs, judgments, and scripts to reproduce our results and serve as tools for such evaluation. Our code is available at http://gitlab.com/quantellence/research/scientific-recall-bench , and the datasets are available at http://huggingface.co/datasets/quantellence/srb-data .
发表机构
- Quantellence Research(Quantellence研究公司)
- St. Petersburg Department of the Steklov Institute of Mathematics(斯捷克洛夫数学研究所圣彼得堡分部)
- St. Petersburg State University(圣彼得堡国立大学)
机构由 AI 辅助整理,请以论文原文为准。