发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多方对话中消息归因与状态重建瓶颈,提出双轨记忆框架 SpeakerMem-R1,结合逐字与结构化记忆及人物/群体视图,在多个基准上取得最优结果。
AI 中文摘要
多方场景中的长期对话记忆不仅需要从长期对话中检索相关内容,还必须区分谁说了什么、每句话涉及谁、个体之间如何相互感知、群体共享哪些信息,以及状态如何随时间变化。近期对多方对话基准的研究表明,现有的通用大语言模型记忆系统往往会丢失人物和群体关系,或难以整合分布在成员、群体和时间中的线索。这些问题共同揭示了两个核心瓶颈:多方对话中的消息归因和关系理解,以及从交错历史中进行状态重建。为解决这两个问题,我们提出了 SpeakerMem-R1:其双轨记忆存储带有说话者标签的逐字消息和派生的状态,这些状态按人物级和群体级视图组织,然后在查询时通过实体、事件和时间结合两条轨道的证据。为了在结构化记忆构建过程中减少归因和更新错误,同时支持本地部署,我们使用 SpeakerLevenshtein 和说话者条件 GRPO 训练 Writer-R1。在 GroupMemBench、SocialMemBench 和 EverMemBench 上,SpeakerMem-R1 的二元准确率分别达到 47.9%、69.2% 和 61.9%。在 EverMind-AI 公开的 EverMemBench 排行榜上,我们取得了 62.33% 的成绩,这是最新最先进框架中的最佳报告结果。我们还在全部 1,986 个 LoCoMo 问题上取得了 70.85% 的准确率,并将其用作两人长期对话边界测试。在一项包含 305 个问题的受控评估中,强化学习将 SFT Writer 的平均准确率从 57.38% 提升至 68.20%。我们报告了二元准确率和 token-F1,消融实验表明,在标准化评估接口下,逐字轨道和结构化轨道以及人物级和群体级视图是互补的。
英文摘要
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed across members, groups, and time. Together, these issues reveal two core bottlenecks: message attribution and relational understanding in multi-party dialogue, and state reconstruction from interleaved histories. To address both, we propose $\textbf{SpeakerMem-R1}$: its dual-track memory stores speaker-labeled verbatim messages and derived states organized into person-level and group-level views, then combines evidence from both tracks by entity, event, and time at query time. To reduce attribution and update errors during structured memory construction while enabling local deployment, we train Writer-R1 with SpeakerLevenshtein and speaker-conditioned GRPO. On GroupMemBench, SocialMemBench, and EverMemBench, SpeakerMem-R1 achieves binary accuracies of 47.9%, 69.2%, and 61.9%, respectively. On the publicly reported EverMemBench leaderboard from EverMind-AI, we achieves 62.33%, the best reported result among the latest state-of-the-art frameworks. It also achieves 70.85% on all 1,986 LoCoMo questions, which we use as a two-person long-term conversation boundary test. In a controlled evaluation of 305 questions, RL raises the SFT Writer's mean accuracy from 57.38% to 68.20%. We report both binary accuracy and token-F1, and ablations show that the verbatim and structured tracks, as well as person-level and group-level views, are complementary under the standardized evaluation interface.
CommentsProject Page: https://2022hpsk.github.io/SpeakerMemR1 , Code: https://github.com/2022hpsk/SpeakerMemR1