通过在线最大成员聚类和原子感知打包实现的紧凑内存大语言模型智能体
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
浏览论文内容
中文总结 AI 辅助
该研究针对长视野LLM部署的紧凑内存需求,提出RSM-full流水线,结合最大成员合并规则与原子感知打包器,在多个基准上实现了质量与令牌成本的优权衡,定义了2k-5k令牌预算下的强帕累托最优。
中文摘要 AI 辅助
许多长视野大语言模型(LLM)部署面临严格的提示预算:随着交互长度增加,延迟、成本和上下文限制使得全上下文提示变得不切实际。关键问题不再仅仅是原始召回率,而是在紧凑内存机制下,哪种内存设计能在质量与令牌之间实现最佳权衡。我们提出了\textbf{RSM-full},这是一种在线聚类内存流水线,旨在实现强的质量-令牌帕累托最优。RSM-full结合了两种设计选择:余弦门控的\textit{最大成员合并}写入规则和原子感知分组上下文打包器。在我们的主要紧凑内存基准AMA-Bench上,在4k预算下,它达到了全上下文质量的83%,同时仅消耗32%的令牌成本;在4次种子平均下,在整个约2.6k至5k的范围内,它击败了最接近的流式聚类基准在线K均值(Online K-Means),提升幅度为3.5至6.0个百分点(p<0.001)。三次种子的消融实验显示,大部分提升来自合并规则(相比在线K均值和匹配τ的DP均值提升5.7个百分点)和分组打包器(相比扁平连接提升5.0个百分点)。该模式在独立的长视野角色记忆基准RealMem上也得到重现:RSM-full相比Budget-RAG提升0.69个百分点(p=0.006),与BM25-RAG相当(配对差值=+0.27个百分点,p=0.47;我们并非在等价检验意义上声称与BM25等价),且显著优于流式原型(Streaming-Proto,提升2.97个百分点)和最接近的2025年智能体内存基准A-MEM(提升1.65个百分点,p<0.001)。在所有基准上,结论一致:在严格预算下,紧凑内存性能主要取决于流式内存的合并方式和检索内容的组装方式。总体而言,RSM-full在约2k至5k提示令牌时最有用,此时它定义了一个强的紧凑内存帕累托最优;在此范围之外,更高令牌的基准仍然更优。
英文摘要
Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-context prompting impractical as interaction length grows. The key question is then not raw recall alone, but which memory design gives the best quality--token trade-off in the compact-memory regime. We present \textbf{RSM-full}, an online clustered-memory pipeline designed for a strong quality--token Pareto point. RSM-full combines two design choices: a cosine-gated \emph{max-member merge} write rule and an atom-aware grouped context packer. On AMA-Bench, our primary compact-memory benchmark, it reaches $83%$ of Full-Context quality at $32%$ of the token cost at a $4$k budget; under four-seed averaging it beats the closest streaming-clustered baseline (Online K-Means) by $+3.5$--$6.0$,pp ($p{<}.001$) across the whole ${\sim}2.6$k--${\sim}5$k regime. Three-seed ablations show most of this gain comes from the merge rule ($+5.7$,pp over Online K-Means and matched-$τ$ DP-means) and the grouped packer ($+5.0$,pp over flat concatenation). The pattern reproduces on RealMem, an independent long-horizon persona-memory benchmark: RSM-full improves on Budget-RAG ($+0.69$,pp, $p{=}.006$), is on par with BM25-RAG (paired $Δ{=}{+}0.27$,pp, $p{=}.47$; we do \emph{not} claim BM25 equivalence in the equivalence-test sense), and significantly outperforms Streaming-Proto ($+2.97$,pp) and the closest reproduced 2025 agentic-memory baseline A-MEM ($+1.65$,pp, $p{<}.001$). Across benchmarks the message is consistent: under tight budgets, compact-memory performance is driven mainly by how streaming memories are merged and how retrieved content is assembled. Overall, RSM-full is most useful when answeroughly $2k$--$5k$ prompt tokens, where itdefines a strong compact-memory Pareto point; higher-token baselines remain stronger outside this regime.