arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体LLM推理中的持久记忆:成本、收益与何时可辨

Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

Hochan Son, Kyungdoe Han, Jaehan Koh, Xiaowu Dai, Wenlu Xu, Guang Cheng

arXiv 2610.07782首次发表:更新:

发表机构

University of California, Los Angeles; University of Wisconsin; HCLTech America(加州大学洛杉矶分校; 威斯康星大学; HCLTech美国公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文测量三层多智能体架构中分解推理与持久记忆的成本收益,发现分解显著降低KV缓存峰值,而持久层无准确率增益,并指出零结果的结构性原因及消融验证条件。

AI 中文摘要

将长上下文推理分解到协作智能体上,将每次调用的活跃KV缓存限制在局部而非总证据量,这在KV缓存内存受限时至关重要。许多此类系统增加一个持久层,用于存储和召回推理轨迹,通常通过报告准确率提升的消融实验来验证。我们在一个三层智能体架构上对两者进行了测量。分解带来的效果:每次查询的峰值KV工作集为14.3 MiB,而单遍和检索增强基线分别为35.5和35.3 MiB。持久层并未带来效果:在八个受控数据集对中,每组n=100,它增加了+0.368 MiB(95%置信区间[+0.167, +0.590])的峰值缓存,且未产生可检测的准确率变化(+0.015,95%置信区间[-0.011, +0.046])。我们认为该零结果是结构性的:单问题基准为每个条目提供其自身的证据并独立评分,且正确性要求在条件之间重置存储的轨迹,因此召回没有可检索的信息性内容。达到这一结论经历了四次测量修正——三次放大了表面上的益处,第四次使该规模的效果看起来可分辨——这些在结果表中均不可见。我们给出了智能体记忆消融必须满足的条件,以及无需了解具体缺陷的检测程序。

英文摘要

Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching it took four measurement corrections -- three inflating the apparent benefit, the fourth making an effect that size look resolvable -- none visible in the results table. We give the conditions an agent-memory ablation must satisfy and detection procedures that need no knowledge of the specific defect.

Comments13 pages, 1 figure. Accepted as a poster at the Machine Learning for Systems Workshop, NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑