多智能体系统的工作负载感知缓存
Workload-Aware Caching for Multi-Agent Systems
浏览论文内容
中文总结 AI 辅助
研究多智能体系统缓存问题,提出结合重新计算成本、DAG 依赖计数和智能体调用频率的工作负载感知逐出策略,经多基准测试评估,该策略能大幅降低延迟,接近无界缓存性能,且与其他优化方法互补。
中文摘要 AI 辅助
多智能体系统将复杂任务分解为专门智能体执行的有向无环图(DAG),为跨查询缓存中间结果创造了天然机会。然而,现有缓存逐出策略基于访问历史统一对待所有缓存条目,忽略了智能体执行环境中独特的结构和工作负载信号。我们提出一种工作负载感知逐出策略,它将重新计算成本、DAG 依赖计数和智能体调用频率这三个信号组合成一个统一的评分函数,在内存受限情况下保留最有价值的条目。在三个跨越不同重用机制的多智能体基准测试中评估,我们的策略相对于未缓存基线最多可将延迟降低 64.7%,比次优有限容量基线平均降低 31.1%的延迟,同时接近无界缓存的性能并保持与所有竞争有限容量方法相当或更高的准确性。我们还表明工作负载感知内容缓存与其他智能体系统优化方法互补,每种技术针对多智能体管道中不同的效率瓶颈。
英文摘要
Multi-agent systems decompose complex tasks into directed acyclic graphs (DAGs) of specialized agent executions, creating natural opportunities for caching intermediate results across queries. However, existing cache eviction policies treat all cached entries uniformly based on access history, ignoring structural and workload signals uniquely available in agentic execution environments. We present a workload-aware eviction policy that combines three signals, namely recomputation cost, DAG dependency count, and agent invocation frequency, into a unified scoring function that retains the most valuable entries under memory constraints. Evaluated across three multi-agent benchmarks spanning diverse reuse regimes, our policy reduces latency by up to 64.7% relative to the uncached baseline and achieves on average a 31.1% latency reduction over the next best finite-capacity baseline, while approaching the performance of an unbounded cache and maintaining accuracy on par with or exceeding all competing finite-capacity methods. We further show that workload-aware content caching is complementary to other agentic system optimization methods, including plan-level caching and parallel agent execution, with each technique targeting a distinct efficiency bottleneck in multi-agent pipelines.