arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38021cs.CLcs.AIcs.IR

可审计的长期记忆:确定性检索链在LongMemEval-S上测得479/475(满分500)

Auditing Long-Term Memory Evaluation: Repeated Judging, Reader Variation, and Negative Controls

Christopher J. Chanhnourack

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在LongMemEval-S上评估可审计长期记忆系统,提出确定性检索链,得分479/475,但未证明优于Chronos High的478,并公开材料供复核。

中文摘要 AI 辅助

我们在LongMemEval-S上评估了一个可审计的长期记忆系统。其检索链采用混合候选检索、交叉编码器重排序、覆盖优先的数据包编译以及确定性推理脚手架;LLM仅用作可替换的最终阅读器。该链将468/470个可回答问题的所有黄金会话放入候选池,并为462/470个问题生成了黄金完整的数据包。通过未固定的CLI别名调用的Claude Opus阅读器,在GPT-4o下两次500问题测试分别得分479/500和475/500。72个可回答的知识更新行使用了实质性修改的评分提示,其在官方文本下的效果尚未测量。这对成绩跨越了Chronos High公布的478/500;阅读器生成、评分提示以及可能的数据版本差异,加上系统内方差,既不能证明优越性也不能证明等价性。同一数据包上的grok-4.6-high阅读器得分476/474,而最大推理努力的智能体变体则退步至461/465。头条通过次数在八个判定翻转行上存在差异。第二位评审员在每次通过中与头条评审员在493/500行(98.6%)上达成一致,并将两次通过均评为472/500;官方评审员在重新评分字节相同的第一次通过答案时也翻转了三个判定。负对照拒绝了一个修复了三个错误草稿但破坏了十一个正确草稿的验证器。所有组件均在相同的500个问题上开发,没有保留验证集或独立的人工裁决;检索和脚手架方法来源及转录衍生的审计被保留;头条阅读器获得了额外的操作员上下文,其完整请求未被保留,且MCP工具可用性尚未解决。我们发布了物化的数据包、脚手架、阅读器输出、评审员判定和对照,以供检查和重新评分。

英文摘要

This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions. Its strongest historical reader lane scores 479 and 475 under an adapted GPT-4o rubric; re-judging the same pass-1 answers changes three labels and yields 478. Fixed-answer knowledge-update re-scoring gives 70/72 under the upstream template and 69/72 under the modified template. Reader lanes span 93 to 479 on fixed packets; paired tests between the two strongest historical lanes establish neither superiority nor equivalence. A different-family reader, configured without client tools or operator files, scores 474, 1.0 percentage point below the headline pass (paired 95% interval [-3.0,+1.0]). Live reader request bodies were not retained. With the same requested reader label, route and judge snapshot, the full package scores 474 versus 454 for baseline sessions, a difference of +4.0 percentage points [95% interval +2.2,+6.0]. Eighteen of the 23 gains, and no losses, occur where baseline packets lacked listed evidence; this post-hoc split does not identify a component effect. In recovered LoCoMo data, token-F1 gains do not survive answer-line extraction. A negative control rejects a verifier that repairs three wrong drafts but breaks eleven correct ones. All questions were used to develop the components; no untouched holdout was evaluated. These findings do not establish a new leaderboard leader or transferable memory advantage. The A/D comparison has one pass per arm, including six reused identical-prompt outcomes, with no pinned reader snapshot; B/C and repeats remain unrun. Original headline requests cannot be reconstructed and stages 1--4 remain closed. Released artifacts support packet inspection and saved-verdict recounting and re-scoring; they do not reconstruct the method.

发表机构

  • Centennial Defense Systems(百年防御系统公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑