arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23568cs.AI

RENDER:控制LLM记忆评估中面向读者的证据

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

Yuan Si, Simeng Han, Daming Li, Jialu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

RENDER是控制LLM记忆评估中面向读者证据的基准,实验显示其相关数据包比原始对话表现更优,效果可迁移,提示需关注评估中的读者面向人工制品。

中文摘要 AI 辅助

记忆与检索增强生成(RAG)评估通常将回答模型的输入视为实现细节,尽管系统可能将相同的历史记录呈现为记忆条目、摘要、类型化记录或原始摘录。我们引入RENDER,这是一种基准控制方法,它在固定对话的同时改变面向读者的人工制品。该方法结合了五级数据包阶梯(用于定位包含答案的内容何时进入输入),以及近似ChatGPT风格条目、LangChain摘要、MemGPT风格类型化记录和原始对话的确定性模板。在500个LongMemEval问题和9个模型上,匹配预算的已解析数据包比按近期截断的原始对话高出42.4-72.6个百分点。在部署风格模板中,每个模型的最佳-最差差值为24.6-48.8个百分点;在主要评分器下,9个模型中有7个的ChatGPT风格条目比原始对话具有更高的点估计值。裁判重评分保留了积极的总体效果,但模型特定的显著性不一。三个在正式账本数据包上得分为0%的模型,从自然语言条目中回答相同事实的比例为45.4-53.4%。该效果在检索噪声下仍然存在,并可迁移到HotpotQA,这表明记忆/RAG评估应报告或控制面向读者的人工制品。

英文摘要

Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt. We introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact. RENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw conversation. On 500 LongMemEval questions and nine models, matched-budget resolved packets beat recency-truncated raw dialogue by 42.4-72.6 points. In deployed-style templates, best-worst spread is 24.6-48.8 points per model; under the primary scorer, ChatGPT-style entries have higher point estimates than raw conversation on 7 of 9 models. Judge rescoring preserves the positive aggregate effect, but model-specific significance is mixed. Three models scoring 0 percent on formal ledger packets answer the same facts from natural-language entries at 45.4-53.4 percent. The effect persists under retrieval noise and transfers to HotpotQA, suggesting that memory/RAG evaluations should report or control the reader-facing artifact.

发表机构

  • University of Waterloo(滑铁卢大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑