AI 中文总结
该研究针对大语言模型智能体记忆系统的延迟与质量矛盾,提出Router-Mem框架,通过证据条件渐进式执行,在降低推理时间的同时保持了优异的任务性能。
AI 中文摘要
大语言模型(LLM)向持续自适应智能发展的过程中,越来越需要能在多轮交互中保存和复用信息的长期记忆机制。现有记忆系统要么压缩并结构化历史以实现高效访问,要么对更广泛的轨迹进行深度研究;前者降低了在线成本,但可能遗漏时间、因果或跨步骤依赖关系,后者提升了证据覆盖范围,但会带来显著的延迟和推理成本。这引发了一个关键问题:能否让记忆系统在保持低在线延迟的同时获得优异的答案质量?我们提出了Router-Mem,一种面向长程智能体记忆的证据条件渐进式执行框架。Router-Mem首先应用共享的低成本检索前缀获取证据,随后一个轻量级充足性路由器预测上下文是否支持提前终止,使其能在推理时做出单token决策;该框架采用证据级监督和原理条件表示蒸馏进行训练。当证据不足时,Router-Mem会复用检索命中结果以扩展记忆块,并进行更深入的分析与聚合。在AMA-Bench和BEAM上的实验表明,Router-Mem分别取得了55.17%和38.77%的分数,与全记忆执行相比,平均推理时间分别减少了27.3%和25.5%。
英文摘要
The continued development of LLMs toward persistent and adaptive intelligence increasingly requires long-term memory mechanisms that preserve and reuse information across interactions. Existing memory systems either compress and structure histories for efficient access or perform deep research over broader trajectories. The former lowers online cost but may omit temporal, causal, or cross-step dependencies, while the latter improves evidence coverage at substantial latency and inference cost. This raises a key question: can a memory system achieve strong answer quality while maintaining low online latency? We introduce Router-Mem, an evidence-conditioned progressive execution framework for long-horizon agent memory. Router-Mem first applies a shared low-cost retrieval prefix to obtain evidence. A lightweight sufficiency router then predicts whether the context supports early termination, which enable a single-token decision at inference time. It is trained with evidence-level supervision and rationale-conditioned representation distillation. When evidence is insufficient, Router-Mem reuses retrieval hits to expand memory blocks and perform deeper analysis and aggregation. Experiments on AMA-Bench and BEAM show that Router-Mem achieves 55.17\% and 38.77\% score while reducing average inference time by 27.3\% and 25.5\% compared with full memory execution.