AI 中文总结
MemoryCPT是一种端到端可训练的智能体记忆框架,通过QAD和QAR阶段优化记忆处理,在LoCoMo和LongMemEval数据集上实现了更优的成本-性能权衡。
AI 中文摘要
长视野大语言模型(LLM)智能体需要记忆系统,能从大量交互历史中恢复有用证据,同时避免向后续模型传递过多上下文。现有记忆流水线常依赖手工启发式规则和重复的LLM调用,这会引入冗余上下文并带来高推理成本。我们提出MemoryCPT,一种端到端可训练的智能体记忆流水线,涵盖离线记忆构建和在线查询条件上下文生成两个阶段。MemoryCPT包含两个阶段:与查询无关的蒸馏(QAD),它利用显式推理轨迹将模块化记忆构建流水线蒸馏为紧凑模型;以及与查询相关的检索与摘要(QAR),它结合 reciprocal rank fusion(RRF)与基于LoRA的摘要器,该摘要器在成本感知奖励下通过Group Relative Policy Optimization(GRPO)训练。我们进一步引入单位成本质量(QPC)来量化每单位推理成本的答案质量。在LoCoMo和LongMemEval上的实验表明,MemoryCPT相较于评估的基线方法改善了成本-性能权衡,而消融和敏感性分析明确了其组件的贡献以及关键设计选择的效果。
英文摘要
Long-horizon LLM agents require memory systems that recover useful evidence from large interaction histories without passing excessive context to downstream models. Existing memory pipelines often rely on hand-crafted heuristics and repeated LLM calls, which can introduce redundant context and high inference cost. We propose MemoryCPT, an end-to-end trainable agent memory pipeline that spans offline memory construction and online query-conditioned context generation. MemoryCPT consists of two stages: Query-agnostic Distillation (QAD), which distills a modular memory-construction pipeline into a compact model using explicit reasoning traces; and Query-aware Retrieval and Summarization (QAR), which combines reciprocal rank fusion (RRF) with a LoRA-based summarizer trained via Group Relative Policy Optimization (GRPO) under a cost-aware reward. We further introduce Quality per Cost (QPC) to quantify answer quality per unit inference cost. Experiments on LoCoMo and LongMemEval show that MemoryCPT improves the cost-performance trade-off over the evaluated baselines, while ablation and sensitivity analyses characterize the contributions of its components and the effects of key design choices.