AMemGym:用于长时对话中助手的交互式记忆基准测试
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
- The Hong Kong University of Science and Technology(香港科技大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
AMemGym通过交互式环境实现长时对话中助手的内存管理优化,揭示现有内存系统性能差距并推动策略自我进化。
AI中文摘要:
在用户与基于大语言模型(LLM)的助手之间进行长时交互时,需要有效的内存管理,但当前的方法在训练和评估内存方面面临挑战。现有的内存基准测试依赖于静态的、非策略性的数据作为上下文,限制了评估的可靠性和可扩展性。为了解决这些差距,我们引入了AMemGym,一个交互式环境,能够实现策略性评估和优化内存驱动的个性化。AMemGym采用结构化数据采样来预定义用户档案、状态依赖的问题以及状态演变轨迹,从而实现低成本生成高质量、评估一致的交互。LLM模拟用户通过角色扮演暴露潜在状态,同时保持结构化状态的一致性。基于结构化数据的综合指标指导助手的评估和优化。广泛的实验揭示了现有内存系统(如RAG、长上下文LLM和代理内存)的性能差距及其相应原因。AMemGym不仅能够有效选择 competing 方法,还可能推动内存管理策略的自我进化。通过将结构化状态演变与自由形式交互联系起来,我们的框架提供了一个可扩展、诊断丰富的环境,以推进对话代理中的内存能力。
英文摘要:
Long-horizon interactions between users and LLM-based assistants necessitate effective memory management, yet current approaches face challenges in training and evaluation of memory. Existing memory benchmarks rely on static, off-policy data as context, limiting evaluation reliability and scalability. To address these gaps, we introduce AMemGym, an interactive environment enabling on-policy evaluation and optimization for memory-driven personalization. AMemGym employs structured data sampling to predefine user profiles, state-dependent questions, and state evolution trajectories, enabling cost-effective generation of high-quality, evaluation-aligned interactions. LLM-simulated users expose latent states through role-play while maintaining structured state consistency. Comprehensive metrics based on structured data guide both assessment and optimization of assistants. Extensive experiments reveal performance gaps in existing memory systems (e.g., RAG, long-context LLMs, and agentic memory) and corresponding reasons. AMemGym not only enables effective selection among competing approaches but also can potentially drive the self-evolution of memory management strategies. By bridging structured state evolution with free-form interactions, our framework provides a scalable, diagnostically rich environment for advancing memory capabilities in conversational agents.