发表机构
University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence(中国科学技术大学; 认知智能国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出HiPS框架,将智能体记忆管理解耦为全局共享基础与用户特定自适应层,通过协同进化优化策略,使记忆增强智能体的响应性能持续优于基线方法。
AI 中文摘要
记忆增强智能体在长期对话中维护紧凑的用户画像,无需处理完整对话历史即可生成个性化且一致的响应。这些用户画像的质量依赖于底层的记忆管理策略:在每一步,智能体必须确定保留、压缩或丢弃的内容。然而,现有方法通常采用训练前确定的静态、一刀切策略。实际上,最优记忆决策本质上是用户特定的,且会随策略优化动态演化。为解决该问题,我们提出HiPS(分层个性化策略,Hierarchical Personalized Strategy)框架,将记忆管理解耦为全局共享基础和用户特定自适应层。具体而言,HiPS采用通用策略(Universal Strategy)从跨角色轨迹中提取共享原则,同时采用角色增量蒸馏(Persona Delta Distillation)为行为偏离一般模式的用户生成定制规则;跨层级规则流(Cross-Level Rule Flow)通过提升经广泛验证的个性化规则、降级矛盾的全局规则来动态校准两者边界。该架构建立了协同进化循环,其中一种机制确保所有策略优化都锚定任务结果。大量实验表明,该方法相较于记忆增强基线取得了持续改进。
英文摘要
Memory-augmented agents maintain compact user profiles throughout extended conversations, enabling personalized and consistent responses without the need to process the entire dialogue history. The quality of these user profiles relies on the underlying memory management strategy: at each step, the agent must determine what to retain, compress, or discard. However, existing methods typically employ a static, one-size-fits-all strategy established before training. In practice, the optimal memory decision is inherently user-specific and dynamically evolves alongside policy optimization. To address this, we propose \textbf{HiPS} (\textbf{Hi}erarchical \textbf{P}ersonalized \textbf{S}trategy), a framework that decouples memory management into a globally shared foundation and a user-specific adaptive tier. Specifically, HiPS employs \textbf{Universal Strategy} to extract shared principles from cross-persona trajectories, alongside \textbf{Persona Delta Distillation} to generate tailored rules for users whose behaviors diverge from general patterns. \textbf{Cross-Level Rule Flow} dynamically calibrates their boundary by promoting broadly validated personal rules and demoting contradicted global ones. The architecture establishes a co-evolution loop where a mechanism guarantees that all strategy refinements are anchored to task outcomes. Extensive experiments demonstrate consistent improvements over memory-augmented baselines.
CommentsEMNLP'2026 Main Conference