RUMBA:俄罗斯用户记忆基准测试
RUMBA: Russian User Memory Benchmark
浏览论文内容
中文总结 AI 辅助
为解决现有大语言模型长期记忆基准测试以英语为中心且无法捕捉相关交互的问题,引入RUMBA基准测试,它有细粒度分类法和统一方法,由带时间戳对话组成,能评估模型并分析其行为,识别记忆机制优缺点。
中文摘要 AI 辅助
大语言模型处理长期记忆的能力愈发关键,但现有基准测试以英语为中心,依赖聚合检索指标,无法捕捉远程上下文、时间信息和推理之间的交互。为解决此问题,我们引入RUMBA(俄罗斯用户记忆基准测试),这是一个用于长期对话记忆的新基准测试,提供了以记忆为中心的问题类型的细粒度分类法和统一方法,涵盖语义类型、会话范围、时间推理和时间表达式的明确性。RUMBA由带时间戳的用户-助手对话组成,其中问答对需要跨会话进行检索、组合和推理。虽然是为俄语设计的,但我们也提供了相同方法下的对齐英语子集。我们评估了当代记忆系统和长上下文模型,并展示了RUMBA如何作为诊断工具来分析跨基准切片的模型行为,识别不同记忆机制的优势和失败模式。
英文摘要
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions. RUMBA consists of timestamped user-assistant dialogues with QA pairs requiring retrieval, combination, and reasoning across sessions. While designed for Russian, we also provide an aligned English subset under the same methodology. We evaluate contemporary memory systems and long-context models, and show how RUMBA serves as a diagnostic tool to analyze model behavior across benchmark slices and identify strengths and failure modes of different memory mechanisms.