AI 中文总结
研究大语言模型智能体长上下文依赖瓶颈问题,提出SF-AMS框架,通过效用驱动机制和复合评分整合信号维护记忆,实验表明该方法在多跳推理等任务中相比基线有显著提升,证明建模记忆重要性对长上下文推理关键。
AI 中文摘要
管理长上下文依赖仍然是大语言模型智能体的主要瓶颈,因为冗余和无关信息会降低多步推理能力。提出了智能体记忆系统的策略性遗忘(SF-AMS)框架,通过对记忆单元的长期重要性进行建模来维护紧凑的高实用性记忆。SF-AMS用效用驱动的生存机制取代静态检索和启发式衰减,从使用冗余和时间信号更新记忆重要性,形成分层记忆结构,过滤噪声并优先处理稳定的实体一致信息。此外,复合重要性评分整合语义和实体级信号以提高检索鲁棒性。在LoCoMo和LongMemEval-s上的实验表明,相对于包括LightMem MemO和A-Mem在内的强大基线有持续提升。在Qwen2.5-7B下的多跳推理中提升最大,在GPT-4o-mini下的时间推理及开放域任务中也有显著提升,证明将记忆重要性建模为动态效用信号对可靠的长上下文推理至关重要。
英文摘要
Managing long-context dependencies remains a primary bottleneck in LLM agents, as redundant and irrelevant information can degrade multi-step reasoning. Strategic Forgetting for Agent Memory Systems (SF-AMS) is proposed as a framework for maintaining compact high-utility memory by modeling the long-term importance of memory units. SF-AMS replaces static retrieval and heuristic decay with a utility-driven survival mechanism that updates memory importance from usage redundancy and temporal signals, inducing a hierarchical memory structure that prioritizes stable entity-consistent information while filtering noise. On top of this, Composite Importance Scoring integrates semantic and entity level signals to improve retrieval robustness. Experiments on LoCoMo and LongMemEval-s show consistent gains over strong state of the art baselines including LightMem MemO and A-Mem. The largest improvement appears in multi-hop reasoning under Qwen2.5-7B where SF-AMS achieves plus 9.65 F1 over the strongest baseline followed by temporal reasoning under GPT-4o-mini plus 6.91 F1 and open-domain tasks plus 6.53 F1 demonstrating strong cross backbone generalization. These results show that modeling memory importance as a dynamic utility signal is critical for reliable long-context reasoning.