发表机构
Called It Inc.(Called It 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MemStrata 通过保留检索主干并添加非重复、带日期、标注说话者的来源片段,在本地 Qwen 3.8 27B 阅读器上实现源感知评分下的高准确率,显著优于仅参考评分和密集检索。
AI 中文摘要
一个充分的对话答案可能不同于简短或不完整的基准参考。为了根据记录的历史来衡量充分性,我们倾向于采用源感知评分,即评判者在评估系统盲答案之前,先对照完整来源检查参考;同时报告原始的仅参考评分。使用本地 Qwen 3.8 27B Q4_K_M 阅读器和 24,000 个令牌的证据上限,MemStrata CL1 在源感知 GPT-5.5 裁决下,于 LongMemEval-S 上得分为 475/500(95.0%),在 LoCoMo 类别 1-4 上得分为 1,400/1,540(90.91%),而相同答案在仅参考评分下分别为 463/500(92.6%)和 1,205/1,540(78.25%)。它保留了检索主干,并添加了非重复、带日期、标注说话者的来源片段。一个使用相同阅读器、证据量约为 4.7 倍的完整历史对照组,在仅参考评分下得分为 464/500,在源感知评分下得分为 470/500(94.0%);这两个差异均不具决定性。在相同预算下仅使用关键词选择的得分为 425,而匹配阅读器的 Letta 分支得分为 438。在 LongMemEval-M 上,其中数据包包含每个历史约 1.6% 的内容,MemStrata CL1 得分为 427/500,损失集中在多会话和时间类问题上。在 300 个 BEAM-1M 问题上,它优于密集检索,得分为 0.738 对 0.706(Wilcoxon p = 0.011)。对未更改请求的相同种子重放改变了 1.5-2.3% 的标签。在相同数据包上,GLM 5.3 flash 在 3 分以内非劣效(462 对 463);Muse Spark 1.3 在 269 个问题上未显示非劣效性。四项预注册干预均未满足其所有注册的进展或可行性标准。签名读取侧工件支持检查,但不能重新生成私有检索管道。源感知评分优于人工裁决的结论尚未确立,且开发暴露、自动化评判依赖以及缺乏留出数据排除了独立复制或排行榜声明。
英文摘要
An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing system-blinded answers; original reference-only grading is reported alongside. With a local Qwen 3.8 27B Q4_K_M reader and a 24,000-token evidence ceiling, MemStrata CL1 scores 475/500 (95.0%) on LongMemEval-S and 1,400/1,540 (90.91%) on LoCoMo categories 1-4 under source-aware GPT-5.5 adjudication, against 463/500 (92.6%) and 1,205/1,540 (78.25%) under reference-only grading of the same answers. It preserves a retrieval backbone and adds nonduplicated, dated, speaker-attributed source spans. A same-reader full-history control with about 4.7 times the evidence scores 464/500 reference-only and 470/500 (94.0%) source-aware; neither difference is decisive. Keyword-only selection at the same budget scores 425, and a matched-reader Letta arm 438. On LongMemEval-M, where the packet holds about 1.6% of each history, MemStrata CL1 scores 427/500, with losses concentrated in multi-session and temporal questions. On 300 BEAM-1M questions it outscores dense retrieval, 0.738 to 0.706 (Wilcoxon p = 0.011). A same-seed replay of unchanged requests changed 1.5-2.3% of labels. On identical packets GLM 5.3 flash is non-inferior within 3 points (462 versus 463); Muse Spark 1.3 did not show non-inferiority on 269 questions. None of four pre-registered interventions met all of its registered advancement or feasibility criteria. Signed read-side artifacts support inspection but do not regenerate the private retrieval pipeline. The superiority of source-aware grading to human adjudication is not established, and development exposure, automated-judge dependence and the absence of held-out data preclude an independent-replication or leaderboard claim.
Comments29 pages, 23 tables, 1 figure. Ancillary files contain per-question grades and reproducible analyses, plus explicitly labelled exports from audited follow-up reports