arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25655cs.CLcs.AI

重构正确的会话片段:超越长上下文的交错式对话记忆评估

Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie

首次发表
浏览论文内容

中文总结 AI 辅助

针对混合主题长对话的记忆难题,研究人员构建了SCALE-QA基准,并提出TSIM方法,在三类LLM后端上均比最强基线提升5.6-17.6个准确率点,实现最优表现。

中文摘要 AI 辅助

与聊天助手的对话日益在单个长会话线程中涵盖多个主题,这对记忆系统构成了挑战。现有的长上下文和记忆基准通常会暴露会话或主题边界,或探查直接的个人记忆问题。这些设置低估了更具挑战性的助手记忆场景:即平坦的混合主题线程,其中系统必须推断早期的哪个会话片段能使后续任务决策有效。我们引入SCALE-QA,这是一个针对平坦未分段线程的约束基础任务QA基准,旨在解决会话片段完整性失败问题。该数据集包含10个领域的3000个经审核的问题,采用确定性四选一多项选择评分,并包含一个确定性运行时构建器;实验使用全部3000个问题(上下文长度为128k)和分层抽样的400个问题诊断集(上下文长度为1M)。SCALE-QA的问题是普通的面向任务的请求,其正确答案取决于对话早期引入的因果相关证据。我们还提出了时间-语义交错记忆重构(TSIM),它将轮次流分割为连贯的会话片段,并通过分层多视图记忆栈对其进行索引,该栈包含确定性的会话片段级摘要和聚类路由视图。实验表明,SCALE-QA对强大的RAG基线和长上下文LLM均构成挑战;在三个开源和专有LLM后端上,TSIM在每种后端设置中均达到最高准确率,比最强的对应基线高出5.6至17.6个准确率点。

英文摘要

Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a flat mixed-topic thread where the system must infer which earlier episode makes a later task decision valid. We introduce SCALE-QA, a constraint-grounded task QA benchmark for flat unsegmented threads targeting episode integrity failure. The dataset contains 3,000 audited questions across 10 domains, uses deterministic four-way multiple-choice grading, and includes a deterministic runtime builder; experiments use all 3,000 questions through 128k and a stratified 400-question diagnostic at 1M. SCALE-QA questions are ordinary task-oriented requests whose correct answer depends on causally related evidence introduced earlier in the conversation. We also propose Temporal-Semantic Interleaved Memory Reconstruction (TSIM), which segments the turn stream into coherent episodes and indexes them through a hierarchical multi-view memory stack with deterministic episode-level summary and cluster-routing views. Experiments show that SCALE-QA challenges strong RAG baselines and long-context LLMs alike; across three open-source and proprietary LLM backends, TSIM achieves the highest accuracy in every backend setting, gaining 5.6-17.6 accuracy points over the strongest corresponding baseline.

发表机构

  • University of California San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑