发表机构
Thakur College of Engineering and Technology(塔库尔工程技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LSREP纵向状态重放协议评估对话记忆,通过ICE v2案例审计揭示其与向量RAG的质量-成本权衡及多会话时间性失败模式。
AI 中文摘要
对话记忆在使用过程中会发生变化,因此仅靠端点问答无法确定持久状态如何累积、老化或纳入修订。我们提出LSREP,一种纵向状态重放评估协议,结合有序重放、显式生命周期调度、重复探测、演进参考答案和机制保真度检查。其架构案例研究是ICE v2,一种本地优先的中间件,具有类型化存储、检索融合和动态上下文预算。私有单用户实例包含1,985轮对话、219个不同探测和52个检查点上的1,211个探测-检查点观察。在三个普通密度数据集上,ICE v2与向量RAG的平均质量差异接近零,同时选择的片段减少32%,但估计提示词元使用量增加6.6%。第四个高密度数据集暴露了无预算基线的灾难性失败。保真度审计限制了归因:程序性检索存在缺陷,多个机制未被使用,图效用未得到证实。在互补的匹配公共诊断中,ICE v2在LongMemEval上明显输给纯向量RAG:仅证据预言机中为50.8%对72.8%,完整S中为43.0%对69.5%。配对差异为-22.0个百分点(95%置信区间[-26.6, -17.4])和-26.5([-31.3, -21.8])。保守弃权(不执行)伴随严重的多会话和时间性失败。ICE在此诊断中使用较少上下文,建立了质量-成本权衡而非优越效率。总之,重放、保真度审计和公共端点测试揭示了架构描述或聚合分数单独无法识别的不同失败模式。
英文摘要
Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation Protocol combining ordered replay, explicit lifecycle schedules, repeated probes, evolving reference answers, and mechanism-fidelity checks. Its architectural case study is ICE v2, a local-first memory middleware with typed stores, retrieval fusion, and dynamic context budgets. The private, single-user instantiation contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints. On three ordinary-density datasets, ICE v2 has a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. A fourth, dense dataset exposes catastrophic failures of the unbudgeted baseline. The fidelity audit limits attribution: procedural retrieval is defective, several mechanisms are unexercised, and graph utility is not established. In a complementary matched public diagnostic, ICE v2 loses decisively to pure vector-RAG on LongMemEval: 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S. Paired differences are -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]). Conservative abstention accompanies severe multi-session and temporal failures. ICE uses less context in this diagnostic, establishing a quality-cost trade-off rather than superior efficiency. Together, replay, fidelity auditing, and public endpoint testing expose distinct failure modes that neither architectural descriptions nor aggregate scores identify alone.
Comments37 pages. Code and evaluation artifacts: https://github.com/Deepnar/ice. The exact system snapshot used for the reported results is preserved in the "v2-paper-eval" tagged release