arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向长期对话的交互式记忆学习

Interactive Memory Learning for Long-Term Conversations

Cai Ke, Jiangyue Yan, Han Zhang, Xin Liu, Zike Yuan, Yue Yu, Hui Wang, Ruifeng Xu

arXiv 2609.17088首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Pengcheng Laboratory(哈尔滨工业大学(深圳); 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长期对话中静态记忆管理无法适应动态用户需求的问题,提出ICML多智能体框架,利用在线强化学习与延迟奖励机制实现记忆策略的交互式学习,显著提升响应质量。

AI 中文摘要

近年来,大型语言模型的进步显著增强了智能体在长期对话建模方面的能力。尽管取得了这些成功,现有方法通常采用静态启发式范式,即信息被被动归档,缺乏自适应的记忆价值评估。因此,这些方法无法自我进化,也无法使其记忆管理与不断变化的用户需求保持一致。为解决这一问题,我们提出了ICML(交互式记忆学习),一个多智能体框架,将记忆机制从被动存档转变为可学习的交互式记忆策略。具体而言,我们首先采用会话合成流程生成专家数据,以促进在未见场景中的快速测试时适应。在此基础上,ICML利用在线强化学习机制,其中规划智能体选择性地编码高价值信息,触发智能体动态检索这些信息以优化响应质量,两个智能体通过持续的交互反馈共同进化。关键在于,两个智能体通过延迟奖励机制同步,该机制将未来反馈传播回早期的存储决策,确保记忆策略与用户期望精确对齐。实验结果表明,ICML显著优于强基线,展现出随着交互积累而持续提升响应质量的独特能力。

英文摘要

Recent advancements in large language models have significantly enhanced the capabilities of agents in modeling long-term conversations. Despite these successes, existing approaches typically adopt a static heuristic paradigm, where information is passively archived without adaptive memory valuation. Consequently, these methods fail to self-evolve or align their memory management with evolving user needs. To address this, we propose ICML (InteraCtive Memory Learning), a multi-agent framework that transforms the memory mechanism from a passive archive into a learnable, interactive memory policy. Specifically, we first employ a session synthesis pipeline to generate expert data, facilitating rapid test-time adaptation in unseen scenarios. Building on this, ICML utilizes an online reinforcement learning mechanism where a Planner agent selectively encodes high-value information and a Trigger agent dynamically retrieves it to optimize response quality, whereby the two agents co-evolve through continuous interaction feedback. Crucially, both agents are synchronized through a delayed reward mechanism that propagates future feedback back to earlier storage decisions, ensuring memory policies are precisely aligned with user expectations. Experimental results demonstrate that ICML significantly outperforms strong baselines, exhibiting the unique capability to continuously improve response quality as interactions accumulate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑