AI 中文总结
针对长期语言模型智能体检索器不积累经验的问题,提出EARM框架,结合观测与估计评分重排序,提升答案准确率并降低LLM重排序推理开销。
AI 中文摘要
长期语言模型智能体在交互过程中积累记忆,但其检索器通常不会积累检索经验。语义检索效率高,但嵌入相似度并不总能反映某条记忆是否包含与当前查询相关的证据。大型语言模型(LLM)重排序器能提供更强的查询条件相关性评分,但无状态重排序会反复对大量候选池评分,且每次查询后丢弃这些评分。我们提出EARM,一种经验摊销重排序框架,将先前获取的LLM相关性评分视为可复用的检索经验。EARM在在线矩阵中存储稀疏的查询-记忆相关性评分,通过因果矩阵补全学习其共享结构,再将少量新观测到的评分与估计评分结合,对剩余候选进行重排序。随着经验积累,评分预算会降低,使LLM重排序从每次查询的重复开销转变为智能体生命周期内习得的检索能力。在长期对话记忆上的实验显示,混合观测与估计的重排序使答案准确率较语义检索提升高达6.62%,且仅17.5%的候选获得直接LLM相关性评分时仍保持有效,从而大幅降低LLM重排序的推理开销。这些结果推动了对智能体记忆的更广泛认知:长寿命智能体不仅应记住过去的内容,还应记住这些内容如何被证明对检索有用。
英文摘要
Long-term language-model agents accumulate memories across interactions, but their retrievers typically do not accumulate retrieval experience. Semantic retrieval is efficient, but embedding similarity does not always reflect whether a memory contains evidence relevant to the current query. Large language model (LLM) rerankers provide stronger query-conditioned relevance scores, yet stateless reranking repeatedly scores a large candidate pool and discards these scores after each query. We introduce EARM, an experience-amortized reranking framework that treats previously acquired LLM relevance scores as reusable retrieval experience. EARM stores sparse query--memory relevance scores in an online matrix, learns their shared structure through causal matrix completion, and combines a small set of newly observed scores with estimated scores to rerank the remaining candidates. The scoring budget decreases as experience accumulates, changing LLM reranking from a repeated per-query expense into a retrieval capability learned over an agent's lifetime. Experiments on long-term conversational memory show that mixed observed-and-estimated reranking improves answer accuracy over semantic retrieval by up to 6.62% and remains effective when only 17.5% of candidates receive direct LLM relevance scores, thereby substantially reducing the inference overhead of LLM reranking. These results motivate a broader view of agent memory: a long-lived agent should remember not only past content, but also how that content has proved useful for retrieval.